Polynucleotide barcode generation
Patent Information
- Application Number
- JP2020206665
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2013-07-10
- Filing Date
- 2020-12-14
- Publication Date
- 2025-06-02
- Estimated Expiration
- 2034-02-07
AI Technical Summary
The high cost of synthesizing polynucleotide barcodes for next-generation sequencing is prohibitive due to the expense per base and the need for diverse sequences, making it impractical for large barcode libraries.
A method for generating polynucleotide barcode libraries involves synthesizing barcode sequences, compartmentalizing them into partitions, amplifying, and isolating them using techniques like emulsion PCR, allowing for the creation of large libraries with diverse barcode sequences at reduced cost.
This approach enables the production of large libraries with diverse barcode sequences efficiently, reducing synthesis costs and enhancing the applicability of polynucleotide barcodes in sequencing applications.
Smart Images

Figure 00000057_0000 
Figure 00000058_0000 
Figure 00000059_0000
Abstract
Description
[Technology Field]
[0001] cross reference This application claims the interests of U.S. Provisional Patent Application No. 61 / 762,435 filed on 8 February 2013, U.S. Provisional Patent Application No. 61 / 800,223 filed on 15 March 2013, U.S. Provisional Patent Application No. 61 / 840,403 filed on 27 June 2013, and U.S. Provisional Patent Application No. 61 / 844,804 filed on 10 July 2013, all of which are incorporated herein by reference for all purposes. [Background technology]
[0002] Polynucleotide barcodes are useful in many applications, including next-generation sequencing technologies. Such barcodes typically contain unique identifier sequences and can be extremely expensive to manufacture in sufficient diversity and scale. The cost of synthesizing a single polynucleotide barcode is a function of the cost per base during synthesis and the length of the polynucleotide. Therefore, the cost of synthesizing multiple barcodes, each with a different sequence, is equal to the cost per base multiplied by the number of bases per molecule and the number of molecules in the multiple barcodes. Currently, synthesizing a DNA sequence costs approximately $0.10 per base. For barcode libraries containing tens of thousands to tens of millions of barcodes, this cost is exorbitant. Therefore, improved methods for generating barcode libraries are greatly needed. [Prior art documents] [Patent Documents]
[0003] [Patent Document 1] U.S. Patent Application Publication No. 2012 / 0211084 [Patent Document 2] U.S. Patent Application Publication No. 2012 / 0132288 [Non-patent literature]
[0004] [Non-Patent Document 1] Genome Analysis: A Laboratory Manual Series (Vols. I-IV)(Cold Spring Harbor Laboratory Press) [Non-licensed document 2] Using Antibodies: A Laboratory Manual (Cold Spring Harbor Laboratory Press) [Non-licensed document 3] Cells: A Laboratory Manual (Cold Spring Harbor Laboratory Press)
Non-licensed Document 4
Non-licensed Document 5
Non-licensed Document 6
Non-licensed Document 7
Non-licensed Document 8
Non-licensed literature 9
Non-licensed literature 10
Non-licensed Document 11
[0005] This disclosure provides methods, compositions, systems, and kits for generating polynucleotide barcodes, as well as uses for such polynucleotide barcodes. Such polynucleotide barcodes can be used for any suitable application. [Means for solving the problem]
[0006] Aspects of the present disclosure provide a library comprising one or more polynucleotides, each of which comprises a barcode sequence, the polynucleotides arranged in one or more partitions, and the library comprising at least about 1,000 different barcode sequences.
[0007] In some cases, the barcode sequence is at least about 5 nucleotides long. Alternatively, the barcode sequence may be a random polynucleotide sequence.
[0008] Furthermore, the partitions may contain, on average, about 1 polynucleotide, about 0.5 polynucleotides, or about 0.1 polynucleotides. The partitions may be droplets, capsules, wells, or beads.
[0009] Furthermore, the library may include at least approximately 10,000 different barcode sequences, at least approximately 100,000 different barcode sequences, at least approximately 500,000 different barcode sequences, at least approximately 1,000,000 different barcode sequences, at least approximately 2,500,000 different barcode sequences, at least approximately 5,000,000 different barcode sequences, at least approximately 10,000,000 different barcode sequences, at least approximately 25,000,000 different barcode sequences, at least approximately 50,000,000 different barcode sequences, or at least approximately 100,000,000 different barcode sequences.
[0010] In some cases, a partition may contain multiple copies of the same polynucleotide.
[0011] Furthermore, each polynucleotide may include a sequence selected from the group consisting of an immobilization sequence, an annealing sequence for sequencing primers, and a sequence suitable for ligation with the target polynucleotide.
[0012] In some cases, each polynucleotide is a MALBAC primer.
[0013] Another aspect of the present disclosure provides a method for synthesizing a library of polynucleotides containing barcode sequences, comprising: a) synthesizing a plurality of polynucleotides containing barcode sequences; b) separating the polynucleotides into a plurality of partitions to produce partitioned polynucleotides; c) amplifying the partitioned polynucleotides to produce amplified polynucleotides; and d) isolating the partitions containing the amplified polynucleotides. In some cases, the synthesis step includes incorporating a mixture of adenine, thymine, guanine, and cytosine into the coupling reaction.
[0014] Furthermore, the separation may include a step of producing diluted polynucleotides by performing limiting dilution. In some cases, the separation further includes a step of partitioning the diluted polynucleotides.
[0015] Furthermore, amplification can be carried out by a method selected from the group consisting of polymerase chain reaction, asymmetric polymerase chain reaction, emulsion PCR (ePCR), ePCR with the use of beads, ePCR with the use of hydrogels, amplification cycle based on multiple annealing and looping (MALBAC), single-primer isothermal amplification, and combinations thereof. In some cases, amplification may be carried out using RNA primers and may include a step of exposing the amplified polynucleotides to RNAase H.
[0016] In some cases, each of the polynucleotides containing the barcode sequence is a MALBAC primer.
[0017] In some cases, isolation can be performed by flow-assisted sorting.
[0018] Furthermore, a hairpin structure can be formed from polynucleotides selected from the group consisting of polynucleotides containing barcode sequences and amplified polynucleotides. In some cases, the method may further include the step of cleaving the hairpin structure within a non-annealing region.
[0019] Furthermore, polynucleotides selected from the group consisting of the barcode sequence-containing polynucleotide, the segmented polynucleotide, and the amplified polynucleotide may be bound to the beads.
[0020] The method may further include the step of annealing the amplified polynucleotide with a partially complementary sequence. The partially complementary sequence may include a barcode sequence.
[0021] The method may further include the step of conjugating at least one of the amplified polynucleotides to a target sequence. The target sequence may be fragmented. In some cases, the target sequence is fragmented by a method selected from the group consisting of mechanical shearing and enzymatic treatment. Mechanical shearing can be induced by ultrasound. In some cases, the enzyme is selected from the group consisting of restriction enzymes, fragmentases, and transposases. Furthermore, conjugation can be carried out by a method selected from the group consisting of ligation and amplification.
[0022] In some cases, the amplification is carried out using a MALBAC primer, thereby producing a MALBAC amplification product. In some cases, the MALBAC primer contains the amplified polynucleotide. In some cases, the MALBAC primer contains a polynucleotide other than the amplified polynucleotide. In such cases, the method may further include the step of conjugating the MALBAC amplification product to the amplified polynucleotide.
[0023] Furthermore, each partition may, on average, contain approximately 1 polynucleotide containing a barcode sequence, 0.5 polynucleotides containing a barcode sequence, or 0.1 polynucleotides containing a barcode sequence. Additionally, the partitions may be selected from the group consisting of droplets, capsules, and wells.
[0024] In some cases, the library may include at least approximately 1,000 different barcode sequences, at least approximately 10,000 different barcode sequences, at least approximately 100,000 different barcode sequences, at least approximately 500,000 different barcode sequences, at least approximately 1,000,000 different barcode sequences, at least approximately 2,500,000 different barcode sequences, at least approximately 5,000,000 different barcode sequences, at least approximately 10,000,000 different barcode sequences, at least approximately 25,000,000 different barcode sequences, at least approximately 50,000,000 different barcode sequences, or at least approximately 100,000,000 different barcode sequences.
[0025] In some cases, a partition contains multiple copies of the same polynucleotide, including the barcode sequence.
[0026] Furthermore, the polynucleotide containing the barcode sequence may include sequences selected from the group consisting of an immobilization sequence, an annealing sequence for sequencing primers, and sequences suitable for ligation with the target polynucleotide.
[0027] A further aspect of the present disclosure provides a library comprising at least about 1,000 beads, wherein each of the at least about 1,000 beads comprises a different barcode sequence. In some cases, the different barcode sequences may be contained within a polynucleotide comprising an immobilization sequence and / or an annealing sequence for sequencing primers. In some cases, the different barcode sequences may be at least about 5 nucleotides long or at least about 10 nucleotides long. In some cases, the different barcode sequences may be random polynucleotide sequences or generated by combination.
[0028] Furthermore, each of the 1,000 beads may contain multiple copies of different barcode sequences. For example, each of the 1,000 beads may contain at least approximately 100,000 copies, at least approximately 1,000,000 copies, or at least approximately 10,000,000 copies of different barcode sequences. In some cases, the library may further contain two or more beads containing the same barcode sequence. In some cases, at least two beads of the 1,000 beads may contain the same barcode sequence. Furthermore, at least approximately 1,000 beads may contain at least approximately 10,000 beads, or at least approximately 100,000 beads.
[0029] The library may also contain at least approximately 1,000, at least approximately 10,000, at least approximately 100,000, at least approximately 1,000,000, at least approximately 2,500,000, at least approximately 5,000,000, at least approximately 10,000,000, at least approximately 25,000,000, at least approximately 50,000,000, or at least approximately 100,000,000 different barcode sequences.
[0030] In some cases, at least approximately 1,000 beads can be distributed across multiple partitions. In some cases, the partitions may be droplets of emulsion. In some cases, each bead of the 1,000 beads may be contained within a different partition. In some cases, the different partitions may be droplets of emulsion. In some cases, two or more beads of the 1,000 beads may be contained within different partitions. In some cases, the different partitions may be droplets of emulsion. In some cases, the 1,000 beads may be hydrogel beads.
[0031] Further aspects of this disclosure provide the use of libraries, compositions, methods, apparatus, or kits described herein in the following applications: species segregation, oligonucleotide segregation, species stimulant selective release from partitions, carrying out reactions in partitions (e.g., ligation and amplification reactions), carrying out nucleic acid synthesis reactions, nucleic acid barcoding, preparation of polynucleotides for sequencing, polynucleotide sequencing, mutation detection, neurological disorder diagnosis, diabetes diagnosis, fetal aneuploidy diagnosis, cancer mutation detection and forensic medicine, disease detection, medical diagnosis, low-input nucleic acid application, circulating tumor cell (CTC) sequencing, polynucleotide phase formation, sequencing of polynucleotides from a small number of cells, gene expression analysis, polynucleotide segregation from cells, or combinations thereof.
[0032] Embedding by reference All publications, patents, and patent applications described herein are incorporated herein by reference to the same extent as if each individual publication, patent, or patent application were specifically and individually indicated as being incorporated by reference.
[0033] Novel features of the methods, compositions, systems, and apparatus of this disclosure will be described in particular, along with the details of the appended claims. A better understanding of the features and advantages of this disclosure will be obtained from the following detailed description illustrating exemplary embodiments in which the principles of the methods, compositions, systems, and apparatus of this disclosure are used, and from reference to the appended drawings. [Brief explanation of the drawing]
[0034] [Figure 1] This is a schematic diagram illustrating an example of a fork-type adapter. [Figure 2] This diagram schematically shows an example of the arrangement of barcode areas. [Figure 3] This figure shows an exemplary sequence of two fork-shaped adapters ligated to the opposite end of a target polynucleotide. [Figure 4] This figure schematically illustrates an example of a method used to produce the fork-type adapter described in Example 1. [Figure 5] This figure schematically shows an example of a capsule-in-a-capsule as described in Example 2. [Figure 6] This figure schematically shows an example of a capsule-in-a-capsule as described in Example 3. [Figure 7] This figure schematically shows examples of products (or intermediates) that can be produced according to the method described in Example 4. [Figure 8a] This figure shows an example of the arrangement described in Example 4. [Figure 8b] This figure shows an example of the arrangement described in Example 4. [Figure 8c] This figure shows an example of the arrangement described in Example 4. [Figure 9a] This figure shows an example of the arrangement described in Example 5. [Figure 9b] This figure shows an example of the arrangement described in Example 5. [Figure 9c] This figure shows an example of the arrangement described in Example 5. [Figure 9d] This figure shows an example of the arrangement described in Example 5. [Figure 9e] This figure shows an example of the arrangement described in Example 5. [Figure 9f] This figure shows an example of the arrangement described in Example 5. [Figure 9g] This figure shows an example of the arrangement described in Example 5. [Figure 9h] This figure shows an example of the arrangement described in Example 5. [Figure 9i] This figure shows an example of the arrangement described in Example 5. [Figure 9j] This figure shows an example of the arrangement described in Example 5. [Figure 10a] This figure shows an example of the arrangement described in Example 6. [Figure 10b] This figure shows an example of the arrangement described in Example 6. [Figure 10c] This figure shows an example of the arrangement described in Example 6. [Figure 10d] This figure shows an example of the arrangement described in Example 6. [Figure 10e] This figure shows an example of the arrangement described in Example 6. [Figure 11a] This figure schematically shows the method and structure described in Example 7. [Figure 11b] This figure schematically shows the method and structure described in Example 7. [Figure 11c] This figure schematically shows the method and structure described in Example 7. [Figure 11d] This figure schematically shows the method and structure described in Example 7. [Figure 12] This diagram schematically illustrates capsule production using an exemplary flow-focusing method. [Figure 13] This diagram schematically illustrates the production of capsule-in-capsule via an exemplary flow-focusing method. [Figure 14a] This figure schematically shows the method and structure described in Example 8. [Figure 14b] This figure schematically shows the method and structure described in Example 8. [Figure 14c] This figure schematically shows the method and structure described in Example 8. [Figure 14d] This figure schematically shows the method and structure described in Example 8. [Figure 14e] This figure schematically shows the method and structure described in Example 8. [Figure 15a] This figure schematically shows the method and structure described in Example 9. [Figure 15b] This figure schematically shows the method and structure described in Example 9. [Figure 15c] This figure schematically shows the method and structure described in Example 9. [Figure 15d] This figure schematically shows the method and structure described in Example 9. [Figure 15e] This figure schematically shows the method and structure described in Example 9. [Figure 16] This figure schematically shows the method and structure described in Example 10. [Figure 17] This diagram schematically shows the capsule-in-capsule described in Example 11. [Figure 18] This figure schematically shows the capsule-in-capsule described in Example 12. [Figure 19a] This figure shows an example of the arrangement described in Example 13. [Figure 19b] This figure shows an example of the arrangement described in Example 13. [Figure 19c] This figure shows an example of the arrangement described in Example 13. [Figure 19d] This figure shows an example of the arrangement described in Example 13. [Figure 19e] This figure shows an example of the arrangement described in Example 13. [Figure 19f] This figure shows an example of the method and structure described in Example 13. [Figure 20] This figure schematically shows the capsule-in-capsule described in Example 14. [Figure 21a] This figure schematically shows the method and structure described in Example 15. [Figure 21b] This figure schematically shows the method and structure described in Example 15. [Figure 21c] This figure schematically shows the method and structure described in Example 15. [Figure 22] This diagram schematically shows the capsule-in-capsule described in Example 16. [Figure 23] This diagram schematically shows the capsule-in-capsule described in Example 17. [Modes for carrying out the invention]
[0035] Although various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided only as examples. Those skilled in the art can make several variations, changes, and substitutions without departing from the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein can be used.
[0036] This disclosure provides methods, compositions, systems, and kits for generating polynucleotide barcodes and for using such polynucleotide barcodes. Such polynucleotide barcodes can be used in any suitable application. In some cases, the polynucleotide barcodes provided in this disclosure can be used in next-generation sequencing reactions. Next-generation sequencing reactions include whole-genome sequencing, detection of specific sequences such as single nucleotide polymorphisms (SNPs) and other mutations, detection of nucleic acid (e.g., deoxyribonucleic acid) insertions, and detection of nucleic acid deletions.
[0037] The use of the methods, compositions, systems, and kits described herein may include any prior art of organic chemistry, polymer technology, microfluidics, molecular biology, recombinant technology, cell biology, biochemistry, and immunology, unless otherwise noted. Such prior art includes well and microwell construction, capsule generation, emulsion generation, spotting, microfluidic device construction, polymer chemistry, restriction digestion, ligation, cloning, polynucleotide sequencing, and polynucleotide sequence assembly. Specific non-limiting examples of preferred techniques are described throughout this disclosure. However, equivalent procedures may also be used. Descriptions of certain techniques can be found in standard laboratory manuals such as Genome Analysis: A Laboratory Manual Series (Vols. I-IV), Using Antibodies: A Laboratory Manual, Cells: A Laboratory Manual, PCR Primer: A Laboratory Manual, and Molecular Cloning: A Laboratory Manual (all from Cold Spring Harbor Laboratory Press), as well as in Oligonucleotide Synthesis: A Practical Approach, 1984, IRL Press London, all of which are incorporated herein by reference for all purposes.
[0038] I. Definition The terms used herein are intended to describe only specific embodiments and are not intended to limit them.
[0039] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless otherwise explicitly indicated in the text. Furthermore, to the extent that the terms "including," "includes," "having," "has," "with," "such as," or variations thereof are used in either the specification or / or claims, such terms are intended to be comprehensive, not restrictive, but in a manner similar to that of "comprising."
[0040] As used herein, the term "about" generally refers to a range that is 15% higher or lower than a given number within the context of its specific use. For example, "about 10" includes the range of 8.5 to 11.5.
[0041] As used herein, the term "barcode" generally refers to a label that can be attached to an analyte to convey information about the analyte. For example, a barcode may be a polynucleotide sequence attached to a fragment of a target polynucleotide contained within a particular partition. This barcode can then be sequenced using the fragment of the target polynucleotide. The presence of the same barcode on multiple sequences can provide information about the origin of the sequence. For example, a barcode may indicate that the sequence originated from a particular partition and / or a proximal region of the genome. This can be particularly useful in sequence assembly when several partitions are pooled before sequencing.
[0042] In this specification, the term "bp" generally refers to an abbreviation for "base pair".
[0043] As used herein, the term "microwell" generally refers to a well with a volume of less than 1 mL. Microwells can be manufactured in various volumes depending on the application. For example, microwells can be manufactured in a size appropriate to accommodate any partition volume described herein.
[0044] As used herein, the term “partition” may be a verb or a noun. When used as a verb (e.g., “to partition” or “partitioning”), the term generally refers to the fractionation (e.g., subdivision) of a species or sample (e.g., polynucleotide) between containers that can be used to isolate one fraction (or subdivision) from another. Such containers are referred to using the noun “partition.” Partitioning can be carried out using, for example, microfluidics, dilution, dispensing, etc. A partition may be, for example, a well, a microwell, a hole, a droplet (e.g., a droplet in an emulsion), a continuous phase of an emulsion, a test tube, a spot, a capsule, a bead, a bead surface in a dilution, or any other suitable container for isolating one fraction of a sample from another. A partition may also include another partition.
[0045] As used herein, the terms "polynucleotide" or "nucleic acid" generally refer to molecules containing multiple nucleotides. Examples of polynucleotides include deoxyribonucleic acid, ribonucleic acid, and their synthetic analogs such as peptide nucleic acid.
[0046] As used herein, the term “species” generally refers to any substance that can be used with the methods, compositions, systems, apparatus, and kits of this disclosure. Examples of species include reagents, analytes, cells, chromosomes, tagged molecules or groups of molecules, barcodes, and any sample containing any of these species. Any suitable species can be used, as will be discussed more fully elsewhere in this disclosure.
[0047] II. Polynucleotide barcoding For certain applications, such as polynucleotide sequencing, a unique identifier ("barcode") may be used to identify the origin of the sequence, for example, to assemble a larger sequence from the sequenced fragments. Therefore, it may be desirable to affix a barcode to the polynucleotide fragments before sequencing. The barcode may take various different forms, such as polynucleotide barcodes. Depending on the specific application, the barcode can be attached to the polynucleotide fragments in a reversible or irreversible manner. Furthermore, the barcode may enable the identification and / or quantification of individual polynucleotide fragments during sequencing.
[0048] Barcodes can be filled into partitions so that one or more barcodes are introduced within a particular partition. In some cases, each partition may contain a different set of barcodes. This can be achieved by dispensing barcodes directly into partitions or by placing barcodes within partitions that are contained within other partitions.
[0049] Barcodes can be filled into partitions at an expected or predicted ratio of barcodes per species to be barcoded (e.g., polynucleotide fragments, polynucleotide chains, cells, etc.). In some cases, barcodes are filled into partitions so that approximately 0.0001, 0.001, 0.1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, or 200000 barcodes are filled per species. In some cases, barcodes are filled into partitions so that more than approximately 0.0001, 0.001, 0.1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, or 200000 barcodes are filled per species. In some cases, barcodes are filled into partitions such that each type contains fewer than approximately 0.0001, 0.001, 0.1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, or 200000 barcodes.
[0050] If there are more than one barcode per polynucleotide fragment, such barcodes may be copies of the same barcode or different barcodes. For example, the conjugation process can be designed to conjugate multiple identical barcodes to a single polynucleotide fragment, or to conjugate multiple different barcodes to a polynucleotide fragment.
[0051] The methods provided herein may include the step of filling a partition with reagents necessary for the binding of barcodes to polynucleotide fragments. In the case of a ligation reaction, reagents including restriction enzymes, ligase enzymes, buffers, adapters, barcodes, etc., can be filled into the partition. In the case of barcoding by amplification, reagents including primers, DNA polymerase, dNTPs, buffers, barcodes, etc., can be filled into the partition. In the case of transposon-mediated barcoding (e.g., NEXTERA), reagents including transpososomes (i.e., the terminal complex of transposase and transposon), buffers, etc., can be filled into the partition. In the case of MALBAC-mediated barcoding, reagents including MALBAC primers, buffers, etc., can be filled into the partition. As described throughout this disclosure, these reagents may be filled directly into the partition or through another partition.
[0052] Barcodes can be ligated to polynucleotide fragments using attached or blunt ends. Alternatively, barcoded polynucleotide fragments can be generated by amplifying the polynucleotide fragment using primers containing barcodes. In some cases, barcoded polynucleotide fragments can be generated using MALBAC amplification of the polynucleotide fragment. Primers used for MALBAC may or may not contain barcodes. If the MALBAC primers do not contain barcodes, barcodes can be added to the MALBAC amplification product by other amplification methods, such as PCR. Barcoded polynucleotide fragments can also be generated using transposon-mediated methods. As with any other species considered in this disclosure, these modules can be contained within the same or different partitions, as required by the assay or process.
[0053] In some cases, barcodes can be combinatorially assembled from smaller components designed to be assembled in a modular form. For example, three modules, 1A, 1B, and 1C, can be combinatorially assembled to produce barcode 1ABC. Such combinatorial assemblies can significantly reduce the cost of synthesizing multiple barcodes. For example, a combinatorial system consisting of modules 3A, 3B, and 3C can be produced from just nine modules. * 3 * It is possible to generate a barcode sequence with 3 = 27 possibilities.
[0054] In some cases, as further described elsewhere in this disclosure, barcodes can be combinatorially assembled by mixing two oligonucleotides and hybridizing and annealing them to produce annealed or partially annealed oligonucleotides (e.g., fork-type adapters). These barcodes may contain one or more nucleotide protrusions to facilitate ligation with the polynucleotide fragment to be barcoded. In some cases, the 5' end of the antisense strand can be phosphorylated to ensure double-stranded ligation. Using this technique, different modules can be assembled by mixing, for example, oligonucleotides A and B, A and C, A and D, B and C, B and D, and so on. As described in more detail elsewhere in this disclosure, the annealed oligonucleotides can be synthesized as a single molecule having a hairpin loop and cleaved after ligation to the polynucleotide to be barcoded.
[0055] As described in more detail elsewhere in this disclosure, the binding of polynucleotides to each other may rely on hybridization-compatible overhangs. For example, hybridization of A and T is often used to ensure ligation compatibility between fragments. In some cases, overhangs can be created by processing with enzymes such as Taq polymerase. In some cases, restriction enzymes can be used to produce cleavage products having a single-base 3' overhang, which may be, for example, A or T. Examples of restriction enzymes that remove a single-base 3' overhang include MnII, HphI, Hpy188I, HpyAV, HpyCH4III, MboII, BciVI, BmrI, AhdI, and XcmI. In other cases, restriction enzymes can produce different overhangs (e.g., 5' overhangs, overhangs larger than a single base). Further restriction enzymes that can be used to produce overhangs include BfuCl and Taq. α Examples include I, BbVI, Bccl, BceAl, BcoDI, BsmAI, and BsmFI.
[0056] III. Generation of a Classified Barcode Library In some cases, the disclosure provides methods for generating partitioned barcode libraries and libraries produced according to such methods. In some cases, the methods provided herein combine random synthesis of DNA sequences, separation into partitions, amplification of the separated sequences, and isolation of the amplified separated sequences in order to provide a library of barcodes contained within partitions.
[0057] a. Random synthesis of polynucleotide barcodes In some cases, the methods described herein utilize randomized methods for polynucleotide synthesis, such as randomized methods for DNA synthesis. During randomized DNA synthesis, any combination of A, C, G, and / or T can be added to the coupling step so that each base type in the coupling step is coupled to a subset of the product. When A, C, G, and T are present in equal concentrations, approximately 1 / 4 of the product contains each base. Due to the sequential coupling step and the randomness of the coupling reaction, 4 n (Here, n is the number of bases in the polynucleotide) This allows for the generation of 4 possible sequences. For example, a library of random polynucleotides of length 6 can be 4 6 While a single random sequence has a member diversity of 4,096, a library of length 10 has a member diversity of 1,048,576. Therefore, it is possible to generate very large and complex libraries. These random sequences can be used as barcodes.
[0058] Furthermore, any suitable synthetic base may be used in conjunction with the present invention. In some cases, the bases included in each coupling step can be varied to synthesize a preferred product. For example, the number of bases present in each coupling step may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more. In some cases, the number of bases present in each coupling step may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more. In some cases, the number of bases present in each coupling step may be less than 2, 3, 4, 5, 6, 7, 8, 9, or 10.
[0059] Furthermore, the concentrations of individual bases may be varied to synthesize a preferred product. For example, any base may be present at a concentration of about 0.1, 0.5, 1, 5, or 10 times that of another base. In some cases, any base may be present at a concentration of at least about 0.1, 0.5, 1, 5, or 10 times that of another base. In some cases, any base may be present at a concentration of less than about 0.1, 0.5, 1, 5, or 10 times that of another base.
[0060] The length of the random polynucleotide sequence may be any suitable length depending on the application. In some cases, the length of the random polynucleotide sequence may be 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides. In some cases, the length of the random polynucleotide sequence may be at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides. In some cases, the length of the random polynucleotide sequence may be less than 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides.
[0061] In some cases, the library is defined by the number of members. In some cases, the library may contain about 256, 1024, 4096, 16384, 65536, 262144, 1048576, 4194304, 16777216, 67108864, 268435456, 1073741824, 4294967296, 17179869184, 68719476736, 2.74878 * 10 11 , or 1.09951 * 10 12 members. In some cases, the library may contain at least about 256, 1024, 4096, 16384, 65536, 262144, 1048576, 4194304, 16777216, 67108864, 268435456, 1073741824, 4294967296, 17179869184, 68719476736, 2.74878 * 10 11 , or 1.09951 * 10 12 members. In some cases, the library may contain about 256, 1024, 4096, 16384, 65536, 262144, 1048576, 4194304, 16777216, 67108864, 268435456, 1073741824, 4294967296, 17179869184, 68719476736, 2.74878 * 1011 , or 1.09951 * 10 12 It may include fewer than a certain number of members. In some cases, the library is a barcode library. In some cases, the barcode library may include at least about 1000, 10000, 100000, 1000000, 2500000, 5000000, 10000000, 25000000, 50000000, or 100000000 different barcode sequences.
[0062] The random barcode library may also contain other polynucleotide sequences. In some cases, these other polynucleotide sequences are natural and non-random and include, for example, primer binding sites, annealing sites for the generation of fork-type adapters, immobilization sequences, and regions that enable annealing with the target polynucleotide sequence and thus enable barcoding of the polynucleotide sequence.
[0063] b. Separation of polynucleotides into partitions After synthesizing polynucleotides containing random barcode sequences, the polynucleotides are partitioned into separate compartments to produce a library of partitioned polynucleotides containing barcode sequences. Any suitable separation method and any suitable partition or partition-within-partition configuration can be used.
[0064] In some cases, partitioning is carried out by diluting a mixture of polynucleotides containing a random barcode sequence such that a given volume of dilution contains less than one polynucleotide on average. The given volume of dilution can then be transferred to partitions. Thus, in any given set of partitions, each partition may contain one or zero polynucleotide molecules.
[0065] In some cases, dilution can be carried out so that each partition contains approximately 0.001, 0.01, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, or 2 or more molecules. In some cases, dilution can be carried out so that each partition contains at least approximately 0.001, 0.01, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, or 2 or more molecules. In some cases, dilution can be carried out so that each partition contains approximately 0.001, 0.01, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, or less than 2 molecules.
[0066] In some cases, partitions of approximately 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% contain a certain number of molecules. In some cases, partitions of at least approximately 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% contain a certain number of molecules. In some cases, partitions of less than approximately 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% contain a certain number of molecules.
[0067] In some cases, approximately 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the partitions contain 1 or fewer polynucleotides. In some cases, at least approximately 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the partitions contain 1 or fewer polynucleotides. In some cases, less than approximately 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the partitions contain 1 or fewer polynucleotides.
[0068] In some cases, the partition is a well, microwell, hole, droplet (e.g., a droplet in an emulsion), continuous phase of an emulsion, test tube, spot, capsule, bead surface, or any other suitable container for isolating one fraction of a sample from another. If the partition contains beads, primers for amplification can be attached to the beads. Partitions are described in more detail elsewhere in this disclosure.
[0069] c. Amplification of segmented polynucleotides Next, the segmented polynucleotides are amplified as described above to generate sufficient material for barcoding the target polynucleotide sequence. Any suitable amplification method can be used, such as polymerase chain reaction (PCR), ligase chain reaction (LCR), helicase-dependent amplification, linear post-exponential PCR (LATE-PCR), asymmetric amplification, digital PCR, degenerate oligonucleotide-primer PCR (DOP-PCR), primer extension pre-amplification PCR (PEP-PCR), ligation-mediated PCR, rolling circle amplification, multiple substitution amplification (MDA), and single-primer isothermal amplification (SPIA), emulsion PCR (ePCR), ePCR including the use of beads, ePCR including the use of hydrogels, amplification cycles based on multiple annealing and looping (MALBAC), and combinations thereof. The MALBAC method is described, for example, in Zong et al., Science, 338(6114), pp. 1622-1626 (2012) (the entire work is incorporated herein by reference).
[0070] In some cases, amplification methods that produce single-stranded products (e.g., asymmetric amplification, SPIA, and LATE-PCR) may be preferred. In some cases, amplification methods that produce double-stranded products (e.g., standard PCR) may be preferred. In some cases, the amplification method exponentially amplifies the segmented polynucleotides. In some cases, the amplification method linearly amplifies the segmented polynucleotides. In some cases, the amplification method amplifies the polynucleotides exponentially first, and then linearly. Furthermore, the polynucleotides can be amplified using a single type of amplification, or the amplification can be completed using a series of different types of amplification. For example, ePCR can be combined with further ePCR cycles, or with different types of amplification.
[0071] Amplification is carried out until a suitable amount of polynucleotides containing the barcode is produced. In some cases, amplification can be carried out for 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, or 60 or more cycles. In some cases, amplification can be carried out for at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, or 60 or more cycles. In some cases, amplification can be carried out for fewer than 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, or 60 cycles.
[0072] In some cases, amplification can be carried out until a certain amount of polynucleotide product is produced in each partition. In some cases, amplification is carried out until the amount of polynucleotide product is approximately 10,000,000,000; 5,000,000,000; 1,000,000,000; 500,000,000; 100,000,000; 50,000,000; 10,000,000; 5,000,000; 1,000,000; 500,000; 400,000; 300,000; 200,000; or 100,000 molecules. In some cases, the reduction is carried out until the amount of polynucleotide product is at least approximately 100,000; 200,000; 300,000; 400,000; 500,000; 1,000,000; 5,000,000; 10,000,000; 50,000,000; 100,000,000; 500,000,000; 1,000,000,000; 5,000,000,000; or 10,000,000,000 molecules. In some cases, the reduction is carried out until the amount of polynucleotide product is approximately 10,000,000,000; 5,000,000,000; 1,000,000,000; 500,000,000; 100,000,000; 50,000,000; 10,000,000; 5,000,000; 1,000,000; 500,000; 400,000; 300,000; 200,000; or less than 100,000 molecules.
[0073] d. Isolation of partitions containing amplified sequences. As described above, in some cases, polynucleotides containing barcodes are partitioned such that each partition contains less than one polynucleotide sequence on average. Therefore, in some cases, the partition fractions do not contain polynucleotides and, consequently, cannot contain amplified polynucleotides. Thus, it may be desirable to separate the partitions containing polynucleotides from the partitions that do not contain polynucleotides.
[0074] In some cases, partitions containing polynucleotides are separated from partitions that do not contain polynucleotides using a flow-based sorting method that can identify polynucleotide-containing partitions. In other cases, polynucleotide-containing partitions can be distinguished from those that do not contain polynucleotides using indicators of polynucleotide presence.
[0075] In some cases, nucleic acid staining can be used to identify partitions containing polynucleotides. Exemplary stains include insertion dyes, minor groove binders, major groove binders, external binders, and bis insertion dyes. Specific examples of such dyes include SYBR Green, SYBR Blue, DAPI, propidium iodide, SYBR Gold, ethidium bromide, acridine, proflavin, acridine orange, acrylflavin, fluorocoumarin, ellipticin, daunomycin, chloroquine, zistamycin D, chromomycin, homidium, mitramycin, ruthenium polypyridyl, anthramycin, phenanthridines and acridines, ethidium bromide, propidium iodide, hexidium iodide, dihydroethidium, ethidium homodimer-1 and -2, ethidium monoazide, ACMA, indoles, imidazoles (e.g., Hoechst 33258, Hoechst 33342, Hoechst 34580 and DAPI), acridine orange (this can also be inserted), 7-AAD, actinomycin D, LDS751, hydroxystilvamidine, SYTOX Blue, SYTOX Green, SYTOX Orange, POPO-1, POPO-3, YOYO-1, YOYO-3, TOTO-1, TOTO-3, JOJO-1, LOLO-1, BOBO-1, BOBO-3, PO-PRO-1, PO-PRO-3, BO-PRO-1, BO-PRO-3, TO-PRO-1, TO-PRO-3, TO-PRO-5, JO-PRO-1, LO-PRO-1, YO-PRO-1, YO-PRO-3, PicoGreen, OliGreen, RiboGreen, SYBR Gold, SYBR Green I, SYBR Green II, SYBR DX, SYTO-40, -41, -42, -43, -44, -45 (blue), SYTO-13, -16, -24, -21, -23, -12, -11, -20, -22, -15, -14, -25 Examples include (Green), SYTO-81, -80, -82, -83, -84, -85 (Orange), SYTO-64, -17, -59, -61, -62, -60, and -63 (Red).
[0076] In some cases, isolation methods such as magnetic separation or sedimentation of particles can be used. Such methods may include, for example, the step of binding the polynucleotide to be amplified, a primer corresponding to the polynucleotide to be amplified, and / or the amplified polynucleotide product to beads. In some cases, the binding of the polynucleotide to be amplified, the primer corresponding to the polynucleotide to be amplified, and / or the polynucleotide product to beads may be done via a photo-unstable linker, such as PC-amino-C6. When a photo-unstable linker is used, light can be used to release the bound polynucleotide from the beads. The beads may be, for example, magnetic beads or latex beads. The beads can then be separated, for example, by magnetic separation or sedimentation. Sedimentation of latex particles can be carried out by centrifugation in a liquid denser than the latex, such as glycerol. In some cases, density gradient centrifugation can be used.
[0077] The beads may be of uniform or non-uniform size. In some cases, the diameter of the beads may be approximately 0.001 μm, 0.01 μm, 0.05 μm, 0.1 μm, 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, 150 μm, 200 μm, 300 μm, 400 μm, 500 μm, 600 μm, 700 μm, 800 μm, 900 μm, or 1 mm. The beads may have a diameter of at least about 0.001 μm, 0.01 μm, 0.05 μm, 0.1 μm, 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, 150 μm, 200 μm, 300 μm, 400 μm, 500 μm, 600 μm, 700 μm, 800 μm, 900 μm, or 1 mm. In some cases, the beads may have a diameter of approximately 0.001 μm, 0.01 μm, 0.05 μm, 0.1 μm, 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, 150 μm, 200 μm, 300 μm, 400 μm, 500 μm, 600 μm, 700 μm, 800 μm, 900 μm, or less than 1 mm. In some cases, the beads may have diameters of approximately 0.001 μm to 1 mm, 0.01 μm to 900 μm, 0.1 μm to 600 μm, 100 μm to 200 μm, 100 μm to 300 μm, 100 μm to 400 μm, 100 μm to 500 μm, 100 μm to 600 μm, 20 μm to 50 μm, 150 μm to 200 μm, 150 μm to 300 μm, or 150 μm to 400 μm.
[0078] In some cases, partitions containing polynucleotides can be isolated by using the differential charge between partitions containing polynucleotides and partitions that do not contain polynucleotides, for example, by performing electrophoresis or dielectrophoresis on the partitions.
[0079] In some cases, polynucleotide-containing particles can be identified using selective expansion or contraction of partitions based on osmotic pressure differences. In some examples, polynucleotide-containing partitions can be isolated by flow fractionation, solvent extraction, differential thawing (e.g., using nucleic acid probes), or freezing.
[0080] Isolating partitions containing polynucleotides provides a library of segmented polynucleotide barcodes with significant diversity, while incurring only the cost of a single bulk synthesis.
[0081] IV. Generating an adapter containing a barcode The barcodes described herein may have a variety of structures. In some cases, the barcodes described herein are part of an adapter. Generally, an "adapter" is a structure used to enable the binding of a barcode to a target polynucleotide. The adapter may include, for example, a barcode, a polynucleotide sequence compatible with ligation with the target polynucleotide, and functional sequences such as primer binding sites and immobilization regions.
[0082] In some cases, the adapter is a fork-type adapter. An example of a fork-type adapter is schematically illustrated in Figure 1. Referring to Figure 1, two copies of the fork-type adapter structure 106 are shown opposite the target polynucleotide 105. Each fork-type adapter includes a first immobilization region 101, a second immobilization region 102, a first sequencing primer region 103, a second sequencing primer region 104, and a pair of partially complementary regions (within 103 and 104) that anneal to each other. Using either the sequencing primer region or the immobilization region, the barcoded polynucleotide can be immobilized, for example, on the surface of a bead. The sequencing primer region can be used, for example, as an annealing site for a sequencing primer. In some cases, the protrusions can be designed to allow compatibility with the target sequence. In Figure 1, the pair of annealed polynucleotides 103 and 104 have a 3'-T protrusion, which fits with the 3'-A protrusion on the target polynucleotide 105. The barcode can be incorporated into any suitable portion of the fork-type adapter. After the fork-type adapter containing the barcode is bound to the target sequence 105, the target polynucleotide can be sequenced using the sequencing primer regions 103 and 104. Another example of a fork-type adapter structure is that used in Illumina® library preparations and NEBNext® Multiplex Oligos for Illumina, available from New England Biolabs®. An example of a non-fork-type adapter is disclosed in Merriman et al., Electrophoresis, 33(23) pp. 3397–3417 (2012) (the entire work is incorporated herein by reference).
[0083] Figure 2 shows three schematic diagrammatic examples of the arrangement of barcode areas within the fork-type adapter described in Figure 1. In one example, barcode 205 (BC1) is located within the first immobilization area 201 or between the first immobilization area 201 and the first sequencing primer area 203. In another example, barcode 206 (BC2) is located inside or adjacent to the first sequencing primer area 203. In yet another example, barcode 207 (BC3) is located within the second immobilization area 202 or between the second immobilization area 202 and the second sequencing primer area 204. Figure 2 shows barcodes on both ends of the target sequence, but this is not essential, as for some applications only one barcode per target sequence is sufficient. However, as described elsewhere in this disclosure, more than one barcode can be used per target sequence.
[0084] Figure 3 provides exemplary sequences (SEQ ID NOs. 1 and 22) of two fork-type adapters ligated to the opposite ends of a target polynucleotide (NNN), showing the barcode regions of each fork-type adapter at the sequence level (bold, nucleotides 30-37, 71-77, 81-87, and 122-129). In Figure 3, nucleotides 1-29 represent the immobilization region of the first fork-type adapter, nucleotides 38-70 represent the sequencing primer region of the first fork-type adapter, nucleotides 78-80 (NNN) represent a target polynucleotide of arbitrary length, nucleotides 88-120 represent the sequencing primer region of the second fork-type adapter, and nucleotides 129-153 represent the immobilization region of the second fork-type adapter.
[0085] V. partition a. General characteristics of partitions As described throughout this disclosure, certain methods, compositions, systems, apparatus, and kits of this disclosure may utilize the subdivision (partitioning) of a particular species into separate partitions. The partitions may be, for example, wells, microwells, holes, droplets (e.g., droplets in an emulsion), a continuous phase of an emulsion, a test tube, a spot, a capsule, the surface of a bead, or any other suitable container for isolating one fraction of a sample or species. Partitions may be used to contain the species for further processing. For example, if the species is a polynucleotide analyte, the further processing may include dissection, ligation, and / or barcoding using the species as a reagent. Any number of apparatus, systems, or containers may be used to hold, support, or contain partitions. In some cases, microwell plates may be used to hold, support, or contain partitions. Any suitable microwell plate may be used, for example, a microwell plate having 96, 384, or 1536 wells.
[0086] Each partition may also contain or be contained within any other suitable partition. For example, a well, microwell, hole, bead surface, or tube may contain droplets (e.g., droplets in emulsion), continuous phase in emulsion, spots, capsules, or any other suitable partition. A droplet may contain capsules, beads, or another droplet. A capsule may contain droplets, beads, or another capsule. These descriptions are merely illustrative, and any suitable combinations and pluralities are also assumed. For example, any suitable partition may contain multiple identical or different partitions. In one example, a well or microwell may contain multiple droplets and multiple capsules. In another example, a capsule may contain multiple capsules and multiple droplets. Any combination of partitions is assumed. Table 1 shows non-limiting examples of partitions that can be combined with each other.
[0087] [Table 1]
[0088] Any partition described herein may include multiple partitions. For example, a partition may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 50, 100, 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10000, or 50000 partitions. A partition may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 50, 100, 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10000, or 50000 partitions. In some cases, a partition may contain 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 50, 100, 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10000, or fewer than 50000 partitions. In some cases, each partition may contain 2 to 50, 2 to 20, 2 to 10, or 2 to 5 partitions.
[0089] A partition may contain any suitable species or mixture of species. For example, in some cases, a partition may contain reagents, analytes, samples, cells, and combinations thereof. A partition containing other partitions may contain certain species within the same partition and certain species within different partitions. Species can be distributed among any suitable partitions depending on the needs of a particular process. For example, any partition in Table 1 may contain at least one first species, and any partition in Table 1 may contain at least one second species. In some cases, the first species may be a reagent and the second species may be an analyte.
[0090] In some cases, the species is a polynucleotide isolated from a cell. For example, in some cases, the polynucleotide (e.g., genomic DNA, RNA, etc.) is isolated from the cell using any suitable method (e.g., a commercially available kit). This polynucleotide can be quantified. The quantified polynucleotide can then be partitioned into multiple partitions as described herein. The partitioning of the polynucleotide can be carried out in predetermined inclusions, depending on the quantification of the assay and its requirements. In some cases, all or many of the partitions (e.g., at least 50%, 60%, 70%, 80%, 90%, or 95%) do not contain overlapping polynucleotides such that separate mixtures of non-overlapping fragments are formed across the multiple partitions. The partitioned polynucleotide can then be processed according to any suitable method known in the art or described herein. For example, the partitioned polynucleotide can be fragmented, amplified, barcoded, etc.
[0091] Species can be partitioned using various methods. For example, species can be diluted and dispensed into multiple partitions. The final dilution of the medium containing the species can be performed such that the number of partitions exceeds the number of species. Dilution can also be used before forming an emulsion or capsule, or before spotting the species onto the substrate. The ratio of the number of species to the number of partitions may be about 0.1, 0.5, 1, 2, 4, 8, 10, 20, 50, 100, or 1000. The ratio of the number of species to the number of partitions may be at least about 0.1, 0.5, 1, 2, 4, 8, 10, 20, 50, 100, or 1000. The ratio of the number of species to the number of partitions may be less than about 0.1, 0.5, 1, 2, 4, 8, 10, 20, 50, 100, or 1000. The ratio of the number of species to the number of partitions may be in the range of approximately 0.1–10, 0.5–10, 1–10, 2–10, 10–100, or 100–1000 or more.
[0092] Furthermore, segmentation can also be performed using piezoelectric droplet generation (e.g., Bransky et al., Lab on a Chip, 2009, 9, pp. 516-520) or surface acoustic waves (e.g., Demirci and Montesano, Lab on a Chip, 2007, 7, pp. 1139-1145).
[0093] The number of partitions used may vary depending on the application. For example, the number of partitions may be approximately 5, 10, 50, 100, 250, 500, 750, 1000, 1500, 2000, 2500, 5000, 7500, or 10,000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100,000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1,000,000, 2000000, 3000000, 4000000, 5000000, 10000000, 20000000 or more. The number of partitions may be at least approximately 5, 10, 50, 100, 250, 500, 750, 1000, 1500, 2000, 2500, 5000, 7500, 10,000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100,000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1,000,000, 2000000, 3000000, 4000000, 5000000, 10000000, or 20000000 or more. The number of partitions may be less than approximately 5, 10, 50, 100, 250, 500, 750, 1000, 1500, 2000, 2500, 5000, 7500, 10,000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100,000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1,000,000, 2000000, 3000000, 4000000, 5000000, 10000000, or 20000000. The number of partitions may be approximately 5 to 1,000,000, 5 to 5,000,000, 5 to 1,000,000, 10 to 10,000, 10 to 5,000, 10 to 1,000, 1,000 to 6,000, 1,000 to 5,000, 1,000 to 4,000, 1,000 to 3,000, or 1,000 to 2,000.
[0094] The number of different barcodes or different sets of barcodes to be categorized may vary, for example, depending on the specific barcodes and / or applications to be categorized. Different sets of barcodes may be, for example, sets of identical barcodes where the same barcode is different between each set. Or, different sets of barcodes may be sets of different barcodes where each set is different in the barcodes it contains. For example, approximately 1, 5, 10, 50, 100, 1000, 10000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 90 It is possible to categorize different barcodes or different sets of barcodes with values of 0,000, 1,000,000, 2,000,000, 3,000,000, 4,000,000, 5,000,000, 6,000,000, 7,000,000, 8,000,000, 9,000,000, 1,000,000, 2,000,000, 5,000,000, and 1,000,0000 or more. In some examples, at least about 1, 5, 10, 50, 100, 1000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, It is possible to categorize different barcodes or different sets of barcodes with values of 900,000, 1000000, 2000000, 3000000, 4000000, 5000000, 6000000, 7000000, 8000000, 9000000, 10000000, 20000000, 50000000, and 100000000 or more.In some examples, approximately 1, 5, 10, 50, 100, 1000, 10000, Different barcodes or different sets of barcodes with values less than 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1000000, 2000000, 3000000, 4000000, 5000000, 6000000, 7000000, 8000000, 9000000, 10000000, 20000000, 50000000, or less than 100000000 can be categorized. There are several examples of how barcodes can be categorized into groups of approximately 1-5, 5-10, 10-50, 50-100, 100-1000, 1000-10000, 10000-100000, 10000-1000000, 10000-10000000, or 10000-100000000.
[0095] Barcodes can be segmented at specific densities. For example, barcodes can be segmented so that each partition is approximately 1, 5, 10, 50, 100, 1000, 10000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100000, 200,000, 300,000, 400,000, 500,000, 600,000, It can be categorized to include barcodes of 700,000, 800,000, 900,000, 1000000, 2000000, 3000000, 4000000, 5000000, 6000000, 7000000, 8000000, 9000000, 10000000, 20000000, 50000000, or 100000000. Each partition has at least approximately 1, 5, 10, 50, 100, 1000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200,000, 300,000, 400,000, 500,000, 600,000, 700, The barcodes can be divided so that each partition contains approximately 1, 5, 10, 50, 10 The barcodes can be categorized to include barcodes with values of 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1000000, 2000000, 3000000, 4000000, 5000000, 6000000, 7000000, 8000000, 9000000, 10000000, 20000000, 50000000, or less than 100000000.The barcodes can be divided so that each partition contains approximately 1-5, 5-10, 10-50, 50-100, 100-1000, 1000-10000, 10000-100000, 100000-1000000, 10000-10000000, 10000-10000000, or 10000-100000000 barcodes per partition.
[0096] Barcodes can be segmented so that identical barcodes are divided at a specific density. For example, identical barcodes can be segmented so that each partition is approximately 1, 5, 10, 50, 100, 1000, 10000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100000, 200,000, 300,000, 400,000, 500,000, 600,000, The items can be categorized to contain the same barcode for 700,000, 800,000, 900,000, 1000000, 2000000, 3000000, 4000000, 5000000, 6000000, 7000000, 8000000, 9000000, 10000000, 20000000, 50000000, or 100000000. Each partition has at least approximately 1, 5, 10, 50, 100, 1000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000 barcodes per partition. It is possible to categorize items so that they contain the same barcode for 00, 800,000, 900,000, 1000000, 2000000, 3000000, 4000000, 5000000, 6000000, 7000000, 8000000, 9000000, 10000000, 20000000, 50000000, and 100000000 or more. The barcodes are distributed across each partition, with approximately 1, 5, 10, 50, 100, 1000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, and 8 per partition. The barcodes can be categorized to include identical barcodes with values of 00,000, 900,000, 1000000, 2000000, 3000000, 4000000, 5000000, 6000000, 7000000, 8000000, 9000000, 10000000, 20000000, 50000000, or less than 100000000.The barcodes can be divided so that each partition contains approximately 1-5, 5-10, 10-50, 50-100, 100-1000, 1000-10000, 10000-100000, 100000-1000000, 10000-10000000, 10000-10000000, or 10000-100000000 identical barcodes per partition.
[0097] Barcodes can be segmented so that different barcodes are separated at a specific density. For example, different barcodes can be segmented so that each partition is approximately 1, 5, 10, 50, 100, 1000, 10000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100000, 200,000, 300,000, 400,000, 500,000, 600,000, It can be categorized to include different barcodes such as 700,000, 800,000, 900,000, 1000000, 2000000, 3000000, 4000000, 5000000, 6000000, 7000000, 8000000, 9000000, 10000000, 20000000, 50000000, or 100000000. Each partition has at least approximately 1, 5, 10, 50, 100, 1000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000 barcodes per partition. It is possible to categorize barcodes to include more than 00, 800,000, 900,000, 1000000, 2000000, 3000000, 4000000, 5000000, 6000000, 7000000, 8000000, 9000000, 10000000, 20000000, 50000000, and 100000000 different barcodes. The barcodes are distributed across each partition, with approximately 1, 5, 10, 50, 100, 1000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, and 8 per partition. It is possible to categorize barcodes to include different barcodes such as 00,000, 900,000, 1000000, 2000000, 3000000, 4000000, 5000000, 6000000, 7000000, 8000000, 9000000, 10000000, 20000000, 50000000, or less than 100000000.The barcodes can be divided so that each partition contains approximately 1-5, 5-10, 10-50, 50-100, 100-1000, 1000-10000, 10000-100000, 100000-1000000, 10000-10000000, 10000-10000000, or 10000-100000000 different barcodes per partition.
[0098] The number of partitions used to segment a barcode may vary, for example, depending on the intended use and / or number of different barcodes to be segmented. For example, the number of partitions used to segment a barcode may be approximately 5, 10, 50, 100, 250, 500, 750, 1000, 1500, 2000, 2500, 5000, 7500, or 10,000, 20000, 30000, 40000, 50000, 60000, 70000, 80000. , 90,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 200,000, 300,000, 400,000, 500,000, 100,000,000, 200,000,000 or more. The number of partitions used to segment the barcode may be at least approximately 5, 10, 50, 100, 250, 500, 750, 1000, 1500, 2000, 2500, 5000, 7500, 10,000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 5000000, 600000, 700000, 800000, 900000, 1000000, 20000000, 3000000, 4000000, 5000000, 10000000, or 20000000 or more. The number of partitions used to segment a barcode may be approximately 5, 10, 50, 100, 250, 500, 750, 1000, 1500, 2000, 2500, 5000, 7500, 10,000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 5000000, 600000, 700000, 800000, 900000, 10000000, 2000000, 3000000, 4000000, 5000000, 10000000, or less than 20000000.The number of partitions used to categorize barcodes may be approximately 5 to 1,000,000, 5 to 5,000,000, 5 to 1,000,000, 10 to 10,000, 10 to 5,000, 10 to 1,000, 1,000 to 6,000, 1,000 to 5,000, 1,000 to 4,000, 1,000 to 3,000, or 1,000 to 2,000.
[0099] As described above, different barcodes or different sets of barcodes (for example, each set containing multiple identical or different barcodes) can be partitioned such that each partition contains different barcodes or different sets of barcodes. In some cases, each partition may contain different sets of identical barcodes. When partitioning different sets of identical barcodes, the number of identical barcodes per partition may vary. For example, approximately 100,000 or more different sets of identical barcodes can be partitioned across approximately 100,000 or more different partitions such that each partition contains different sets of identical barcodes. Within each partition, the number of identical barcodes per set of barcodes may be approximately 1,000,000 identical barcodes. In some cases, the number of different sets of barcodes may be equal to or substantially equal to the number of partitions. By combining any suitable number of different barcodes or different barcode sets (including the number of different segmented barcodes or different barcode sets described elsewhere in this specification), the number of barcodes per partition (including the number of barcodes per partition described elsewhere in this specification), and the number of partitions (including the number of partitions described elsewhere in this specification), a diverse library of segmented barcodes containing a large number of barcodes per partition can be generated. Thus, as can be understood, any of the above different numbers of barcodes can be provided together with any of the above barcode densities per partition and in any of the above number of partitions.
[0100] The volume of the partitions may vary depending on the application. For example, the volumes of any partition described herein [e.g., wells, spots, droplets (e.g., in emulsions), and capsules] are approximately 1000 μl, 900 μl, 800 μl, 700 μl, 600 μl, 500 μl, 400 μl, 300 μl, 200 μl, 100 μl, 50 μl, 25 μl, 10 μl, 5 μl, 1 μl, 900 nL, 800 nL, 700 nL, 600 nL, 500 nL, 400 nL, 300 nL, 200 nL, and 100 nL. It may also be 50nL, 25nL, 10nL, 5nL, 2.5nL, 1nL, 900pL, 800pL, 700pL, 600pL, 500pL, 400pL, 300pL, 200pL, 100pL, 50pL, 25pL, 10pL, 5pL, 1pL, 900fL, 800fL, 700fL, 600fL, 500fL, 400fL, 300fL, 200fL, 100fL, 50fL, 25fL, 10fL, 5fL, 1fL, or 0.5fL. The partition volumes are at least approximately 1000 μl, 900 μl, 800 μl, 700 μl, 600 μl, 500 μl, 400 μl, 300 μl, 200 μl, 100 μl, 50 μl, 25 μl, 10 μl, 5 μl, 1 μl, 900 nL, 800 nL, 700 nL, 600 nL, 500 nL, 400 nL, 300 nL, 200 nL, 100 nL, 50 nL, 25 nL, 10 nL, 5 nL, 5 nL, It may also be 2.5nL, 1nL, 900pL, 800pL, 700pL, 600pL, 500pL, 400pL, 300pL, 200pL, 100pL, 50pL, 25pL, 10pL, 5pL, 1pL, 900fL, 800fL, 700fL, 600fL, 500fL, 400fL, 300fL, 200fL, 100fL, 50fL, 25fL, 10fL, 5fL, 1fL, or 0.5fL.The partition capacities are approximately 1000 μl, 900 μl, 800 μl, 700 μl, 600 μl, 500 μl, 400 μl, 300 μl, 200 μl, 100 μl, 50 μl, 25 μl, 10 μl, 5 μl, 1 μl, 900 nL, 800 nL, 700 nL, 600 nL, 500 nL, 400 nL, 300 nL, 200 nL, 100 nL, 50 nL, 25 nL, 10 nL, 5 nL, 5 nL, and 2.5 The volume may be nL, 1nL, 900pL, 800pL, 700pL, 600pL, 500pL, 400pL, 300pL, 200pL, 100pL, 50pL, 25pL, 10pL, 5pL, 1pL, 900fL, 800fL, 700fL, 600fL, 500fL, 400fL, 300fL, 200fL, 100fL, 50fL, 25fL, 10fL, 5fL, 1fL, or less than 0.5fL. The partition volume may be approximately 0.5fL to 5pL, 10pL to 10nL, 10nL to 10μl, 10μl to 100μl, or 100μl to 1mL.
[0101] There may be variability in the volume of liquid in different partitions. More specifically, the volume of different partitions may vary by at least (or more) ±1%, 2%, 3%, 4%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, 500%, or 1000% across the set of partitions. For example, a well (or other partition) may contain a volume of liquid that is at most 80% of the volume of liquid in a second well (or other partition).
[0102] Furthermore, specific species can be targeted to specific partitions. For example, in some cases, capture reagents (e.g., oligonucleotide probes) can be immobilized or positioned within a partition to capture specific species (e.g., polynucleotides). For instance, capture oligonucleotides can be immobilized on the surface of beads to capture species containing oligonucleotides with complementary sequences.
[0103] Furthermore, species can be segmented by specific densities. For example, species can be segmented such that each partition contains approximately 1, 5, 10, 50, 100, 1000, 10000, 100000, or 1,000,000 species per partition. Species can be segmented such that each partition contains at least approximately 1, 5, 10, 50, 100, 1000, 10000, 100000, or 1,000,000 species per partition. Species can be segmented such that each partition contains approximately 1, 5, 10, 50, 100, 1000, 10000, 100000, or less than 1,000,000 species per partition. The species can be divided such that each partition contains approximately 1-5, 5-10, 10-50, 50-100, 100-1000, 1000-10000, 10000-100000, or 100000-1000000 species per partition.
[0104] Species can be partitioned such that at least one partition contains a species that is unique within that partition. This may apply to partitions with approximately 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% or more. This may apply to partitions with at least approximately 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% or less. This may apply to partitions with less than approximately 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%.
[0105] b. Wells as partitions In some cases, wells are used as partitions. Wells may be microwells. Wells may contain a medium containing one or more species. Species may be contained within the wells in various configurations. In one example, species are dispensed directly into the wells. Species dispensed directly into the wells may be covered with a layer that is, for example, soluble, meltable, or permeable. This layer may be, for example, oil, wax, or a film. The layer can be dissolved or melted before or after introducing another species into the well. Wells can be sealed at any point, for example, after the addition of any species, using a sealing layer.
[0106] In one example, reagents for sample processing are dispensed directly into wells and covered with a layer that is soluble, meltable, or permeable. The sample containing the analyte to be processed is introduced on top of the layer. The layer is dissolved or melted, or the analyte (or reagent) diffuses through the layer. The wells are sealed and incubated under appropriate conditions for analyte processing. The processed analyte can then be collected.
[0107] In some cases, a well contains other partitions. A well may contain any suitable partitions, such as another well, a spot, a droplet (e.g., a droplet in an emulsion), a capsule, or a bead. Each partition may exist as a single partition or multiple partitions, and each partition may contain the same or different types.
[0108] In one example, the well contains a capsule containing a reagent for sample processing. The capsule can be filled into the well using a liquid medium or without a liquid medium (e.g., essentially dry). As described elsewhere in this disclosure, the capsule may contain one or more capsules or other partitions. A sample containing the analyte to be processed can be introduced into the well. The well can be sealed and a stimuli applied to induce the release of the capsule contents into the well, bringing the reagent into contact with the analyte to be processed. The well can be incubated under appropriate conditions for processing the analyte. The processed analyte can then be collected. This example describes an embodiment in which the reagent is in a capsule and the analyte is in the well, but the opposite configuration, i.e., the reagent is in the well and the analyte is in a capsule, is also possible.
[0109] In another example, the well contains an emulsion, and the droplets of the emulsion contain capsules containing reagents for sample processing. The sample, containing the analyte to be processed, is contained within the droplets of emulsion. The well is sealed, and a stimuli are applied to induce the release of the capsule contents into the droplets, bringing the reagent into contact with the analyte to be processed. The well is incubated under appropriate conditions for processing the analyte. The processed analyte can then be collected. While this example describes an embodiment where the reagent is in a capsule and the analyte is in a droplet, the opposite configuration, i.e., where the reagent is in a droplet and the analyte is in a capsule, is also possible.
[0110] The wells can be arranged as an array, for example, a microwell array. Based on the dimensions of the individual wells and the size of the substrate, the well array may have various well densities. In some cases, the well density is 10 wells / cm³. 2 50 wells / cm 2 100 wells / cm 2 500 wells / cm 21000 wells / cm 2 5000 wells / cm 2 10,000 wells / cm² 2 50,000 wells / cm² 2 , or 100,000 wells / cm² 2 It may also be the case that the well density is at least 10 wells / cm³. 2 50 wells / cm 2 100 wells / cm 2 500 wells / cm 2 1000 wells / cm 2 5000 wells / cm 2 10,000 wells / cm² 2 50,000 wells / cm² 2 , or 100,000 wells / cm² 2 It may also be the case that the well density is 10 wells / cm³. 2 50 wells / cm 2 100 wells / cm 2 500 wells / cm 2 1000 wells / cm 2 5000 wells / cm 2 10,000 wells / cm² 2 50,000 wells / cm² 2 , or 100,000 wells / cm² 2 It is acceptable to be less than [a certain value].
[0111] c. Spot as a partition In some cases, spots are used as partitions. Spots can be created, for example, by dispensing a substance onto a surface. Seeds may be contained within the spot in various configurations. In one example, seeds are dispensed directly into the spot by incorporating them into the medium used to form the spot. Seeds dispensed directly onto the spot can be covered with a layer that is, for example, soluble, meltable, or permeable. This layer may be, for example, oil, wax, or a film. The layer can be dissolved or melted before or after introducing another seed onto the spot. The spot can be sealed by an overlay at any point, for example, after the addition of any seed.
[0112] In one example, reagents for sample processing are dispensed directly onto a spot, for example, a glass slide, and covered with a layer that is soluble, meltable, or permeable. The sample containing the analyte to be processed is introduced onto the top of the layer. The layer is dissolved or melted, or the analyte (or reagent) diffuses through the layer. The spot is sealed and incubated under appropriate conditions for processing the analyte. The processed analyte can then be collected.
[0113] Spots can also be arranged within wells as described elsewhere in this disclosure (for example, in Table 1). In some cases, multiple spots can be arranged within a well so that the contents of each spot do not mix. Such a configuration may be useful, for example, when it is desirable that the species do not come into contact with each other. In some cases, a well may contain 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more spots. In some cases, a well may contain at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more spots. In some cases, a well may contain 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or fewer than 30 spots. In some cases, a well may contain 2–4, 2–6, 2–8, 4–6, 4–8, 5–10, or 4–12 spots. When adding a substance (e.g., a medium containing the analyte) to the well, the species in the spots may mix. Furthermore, the use of separate spots containing different species (or combinations of species) may also be useful in preventing cross-contamination of the apparatus used to place the spots into the well.
[0114] In some cases, a spot may include other partitions. A spot may include any suitable partition, such as another spot, a droplet (e.g., a droplet in an emulsion), a capsule, or a bead. Each partition may exist as a single partition or multiple partitions, and each partition may contain the same or different types.
[0115] In one example, the spot includes a capsule containing a reagent for sample processing. As described elsewhere in this disclosure, the capsule may contain one or more capsules or other partitions. The sample containing the analyte to be processed is introduced into the spot. The spot is sealed and a stimuli are applied to induce the release of the capsule contents into the spot, bringing the reagent into contact with the analyte to be processed. The spot is incubated under appropriate conditions for processing the analyte. The processed analyte can then be recovered. This example describes an embodiment in which the reagent is in a capsule and the analyte is in the spot, but the opposite configuration, i.e., the reagent is in the spot and the analyte is in a capsule, is also possible.
[0116] In another example, the spot contains an emulsion, and the droplet of the emulsion contains a capsule containing a reagent for sample processing. The sample, containing the analyte to be processed, is contained within the droplet of emulsion. The spot is sealed and a stimuli are applied to induce the release of the capsule contents into the droplet, bringing the reagent into contact with the analyte to be processed. The spot is incubated under appropriate conditions for processing the analyte. The processed analyte can then be recovered. While this example describes an embodiment where the reagent is in a capsule and the analyte is in a droplet, the opposite configuration, i.e., where the reagent is in a droplet and the analyte is in a capsule, is also possible.
[0117] The spots may be of uniform or non-uniform size. In some cases, the diameter of the spots may be approximately 0.1 μm, 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, 150 μm, 200 μm, 300 μm, 400 μm, 500 μm, 600 μm, 700 μm, 800 μm, 900 μm, 1 mm, 2 mm, 5 mm, or 1 cm. The spots may have a diameter of at least approximately 0.1 μm, 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, 150 μm, 200 μm, 300 μm, 400 μm, 500 μm, 600 μm, 700 μm, 800 μm, 900 μm, 1 mm, 1 mm, 2 mm, 5 mm, or 1 cm. In some cases, the spots may have diameters of approximately 0.1 μm, 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, 150 μm, 200 μm, 300 μm, 400 μm, 500 μm, 600 μm, 700 μm, 800 μm, 900 μm, 1 mm, 1 mm, 2 mm, 5 mm, or less than 1 cm. In some cases, the spots may have diameters of approximately 0.1 μm to 1 cm, 100 μm to 1 mm, 100 μm to 500 μm, 100 μm to 600 μm, 150 μm to 300 μm, or 150 μm to 400 μm.
[0118] The spots can be arranged in an array, for example, a spot array. Depending on the dimensions of the individual spots and the size of the substrate, the spot array may have a variety of spot densities. In some cases, the spot density is 10 spots / cm². 2 50 spots / cm 2 100 spots / cm 2 500 spots / cm 2 1000 spots / cm 2 5000 spots / cm 2 10,000 spots / cm 2 50,000 spots / cm 2 , or 100,000 spots / cm 2 This may also be the case. In some cases, the spot density is at least 10 spots / cm². 2 50 spots / cm 2 100 spots / cm 2500 spots / cm 2 1000 spots / cm 2 5000 spots / cm 2 10,000 spots / cm 2 50,000 spots / cm 2 , or 100,000 spots / cm 2 It may also be the case that the spot density is 10 spots / cm². 2 50 spots / cm 2 100 spots / cm 2 500 spots / cm 2 1000 spots / cm 2 5000 spots / cm 2 10,000 spots / cm 2 50,000 spots / cm 2 , or 100,000 spots / cm 2 It is acceptable to be less than [a certain value].
[0119] d. Emulsion as a partition In some cases, droplets in the emulsion are used as partitions. The emulsion can be prepared by any preferred method, such as methods known in the art (see, for example, Weizmann et al., Nature Methods, 2006, 3(7):545-550; Weitz et al., U.S. Patent Application Publication No. 2012 / 0211084). In some cases, aqueous emulsion in carbon fluoride can be used. These emulsions may contain fluorinated surfactants such as oligomer perfluoropolyethers (PFPE) with polyethylene glycol (PEG) (Holtze et al., Lab on a Chip, 2008, 8(10):1632-1639). In some cases, monodisperse emulsions can be formed in a microfluidic flow focusing apparatus (Garstecki et al., Applied Physics Letters, 2004, 85(13):2649-2651).
[0120] The seeds can be contained, for example, in droplets in an emulsion containing a first phase (e.g., oil or water) forming a droplet and a second (continuous) phase (e.g., water or oil). The emulsion may be a single emulsion, for example, water in oil or oil in water emulsion. The emulsion may also be a double emulsion, for example, water in oil or oil in water emulsion. Higher-order emulsions are also possible. The emulsion can be held in any suitable container, such as any suitable partition described in this disclosure.
[0121] In some cases, droplets in the emulsion include other partitions. Droplets in the emulsion may include any suitable partitions, such as other droplets (e.g., droplets in the emulsion), capsules, beads, etc. Each partition may exist as a single partition or multiple partitions, and each partition may contain the same or different species.
[0122] In one example, the droplets in the emulsion contain capsules containing reagents for sample processing. As described elsewhere in this disclosure, the capsules may contain one or more capsules or other partitions. The sample containing the analyte to be processed is placed within the droplet. A stimulus is applied to induce the release of the capsule contents into the droplet, bringing the reagent into contact with the analyte to be processed. The droplet is incubated under appropriate conditions for processing the analyte. The processed analyte can then be recovered. This example describes an embodiment in which the reagent is in a capsule and the analyte is in a droplet, but the opposite configuration, i.e., the reagent is in a droplet and the analyte is in a capsule, is also possible.
[0123] The droplets in the emulsion may be of uniform or non-uniform size. In some cases, the diameter of the droplets in the emulsion may be approximately 0.001 μm, 0.01 μm, 0.05 μm, 0.1 μm, 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, 150 μm, 200 μm, 300 μm, 400 μm, 500 μm, 600 μm, 700 μm, 800 μm, 900 μm, or 1 mm. The droplets may have a diameter of at least about 0.001 μm, 0.01 μm, 0.05 μm, 0.1 μm, 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, 150 μm, 200 μm, 300 μm, 400 μm, 500 μm, 600 μm, 700 μm, 800 μm, 900 μm, or 1 mm. In some cases, the droplets may have diameters of approximately 0.001 μm, 0.01 μm, 0.05 μm, 0.1 μm, 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, 150 μm, 200 μm, 300 μm, 400 μm, 500 μm, 600 μm, 700 μm, 800 μm, 900 μm, or less than 1 mm. In some cases, droplets may have diameters of approximately 0.001 μm to 1 mm, 0.01 μm to 900 μm, 0.1 μm to 600 μm, 100 μm to 200 μm, 100 μm to 300 μm, 100 μm to 400 μm, 100 μm to 500 μm, 100 μm to 600 μm, 150 μm to 200 μm, 150 μm to 300 μm, or 150 μm to 400 μm.
[0124] The droplets in the emulsion may also have a specific density. In some cases, the droplets are less dense than aqueous liquids (e.g., water); in some cases, the droplets are denser than aqueous liquids. In some cases, the droplets are less dense than non-aqueous liquids (e.g., oil); in some cases, the droplets are denser than non-aqueous liquids. The droplets have a density of approximately 0.05 g / cm³. 3 , 0.1 g / cm³ 3 , 0.2 g / cm³ 3 , 0.3g / cm³ 3 , 0.4 g / cm³ 3 , 0.5 g / cm 3 , 0.6 g / cm³ 3 , 0.7 g / cm³ 3 0.8 g / cm³ 3 , 0.81 g / cm³3 , 0.82 g / cm 3 , 0.83 g / cm 3 , 0.84 g / cm 3 , 0.85 g / cm 3 , 0.86 g / cm 3 , 0.87 g / cm 3 , 0.88 g / cm 3 , 0.89 g / cm 3 , 0.90 g / cm 3 , 0.91 g / cm 3 , 0.92 g / cm 3 , 0.93 g / cm 3 , 0.94 g / cm 3 , 0.95 g / cm 3 , 0.96 g / cm 3 , 0.97 g / cm 3 , 0.98 g / cm 3 , 0.99 g / cm 3 , 1.00 g / cm 3 , 1.05 g / cm 3 , 1.1 g / cm 3 , 1.2 g / cm 3 , 1.3 g / cm 3 , 1.4 g / cm 3 , 1.5 g / cm 3 , 1.6 g / cm 3 , 1.7 g / cm 3 , 1.8 g / cm 3 , 1.9 g / cm 3 , 2.0 g / cm 3 , 2.1 g / cm 3 , 2.2 g / cm 3 , 2.3 g / cm 3 , 2.4 g / cm 3 , or 2.5 g / cm 3 may have a density of. The droplet is at least about 0.05 g / cm 3 , 0.1 g / cm 3 , 0.2 g / cm 3 , 0.3 g / cm 3 , 0.4 g / cm 3 , 0.5 g / cm 3 , 0.6 g / cm 3 , 0.7 g / cm 3 , 0.8 g / cm 3 , 0.81 g / cm 3, 0.82 g / cm³ 3 0.83 g / cm³ 3 0.84 g / cm³ 3 , 0.85 g / cm³ 3 , 0.86 g / cm³ 3 0.87 g / cm³ 3 , 0.88 g / cm³ 3 0.89 g / cm³ 3 0.90 g / cm³ 3 , 0.91 g / cm³ 3 , 0.92 g / cm³ 3 0.93 g / cm³ 3 0.94 g / cm³ 3 0.95 g / cm³ 3 , 0.96 g / cm³ 3 0.97 g / cm³ 3 , 0.98 g / cm³ 3 0.99 g / cm³ 3 , 1.00 g / cm³ 3 1.05 g / cm³ 3 , 1.1 g / cm³ 3 , 1.2 g / cm³ 3 1.3 g / cm³ 3 1.4 g / cm³ 3 1.5 g / cm³ 3 1.6 g / cm³ 3 1.7 g / cm³ 3 1.8 g / cm³ 3 1.9 g / cm³ 3 2.0 g / cm³ 3 , 2.1 g / cm³ 3 , 2.2 g / cm³ 3 2.3 g / cm³ 3 2.4 g / cm³ 3 , or 2.5 g / cm³ 3 It may have a density of . Otherwise, the density of the droplet is at most about 0.7 g / cm³. 3 , 0.8 g / cm³ 3 , 0.81 g / cm³ 3 , 0.82 g / cm³ 3 0.83 g / cm³ 3 0.84 g / cm³ 3 , 0.85 g / cm³ 3 , 0.86 g / cm³ 3 , 0.87 g / cm³ 3 0.88 g / cm³ 30.89 g / cm³ 3 0.90 g / cm³ 3 , 0.91 g / cm³ 3 , 0.92 g / cm³ 3 0.93 g / cm³ 3 0.94 g / cm³ 3 0.95 g / cm³ 3 , 0.96 g / cm³ 3 0.97 g / cm³ 3 , 0.98 g / cm³ 3 0.99 g / cm³ 3 , 1.00 g / cm³ 3 1.05 g / cm³ 3 , 1.1 g / cm³ 3 , 1.2 g / cm³ 3 1.3 g / cm³ 3 1.4 g / cm³ 3 1.5 g / cm³ 3 1.6 g / cm³ 3 1.7 g / cm³ 3 1.8 g / cm³ 3 1.9 g / cm³ 3 2.0 g / cm³ 3 , 2.1 g / cm³ 3 , 2.2 g / cm³ 3 2.3 g / cm³ 3 2.4 g / cm³ 3 , or 2.5 g / cm³ 3 This may also be the case. Such density may reflect the density of the capsule in any particular liquid (e.g., aqueous, water, oil, etc.).
[0125] e. Capsule as a partition In some cases, capsules are used as partitions. Capsules can be prepared by any preferred method, including methods known in the art, such as emulsion polymerization (Weitz et al., U.S. Patent Application Publication No. 2012 / 0211084), layer-by-layer assembly with polyelectrolytes, droplet formation, internal phase separation, and flow focusing. Any preferred species can be contained within the capsules. The capsules can be held in any preferred container, such as any preferred partition described in this disclosure.
[0126] In some cases, the capsule contains other partitions. The capsule may contain any suitable partitions, such as another capsule, droplets in an emulsion, or beads. Each partition may exist as a single partition or multiple partitions, and each partition may contain the same or different species.
[0127] In one example, the outer capsule contains an inner capsule. The inner capsule contains a reagent for sample processing. The analyte is encapsulated in a medium between the inner and outer capsules. A stimulus is applied to induce the release of the contents of the inner capsule into the outer capsule, bringing the reagent into contact with the analyte to be processed. The outer capsule is incubated under appropriate conditions for processing the analyte. The processed analyte can then be recovered. This example describes an embodiment in which the reagent is in the inner capsule and the analyte is in the medium between the inner and outer capsules, but the opposite configuration is also possible, i.e., the reagent is in the medium between the inner and outer capsules and the analyte is in the inner capsule.
[0128] Capsules can be pre-formed and filled with reagents by injection. For example, reagents can be introduced into the capsules described herein using the pico-injection method described by Abate et al. (Proc. Natl. Acad. Sci. USA, 2010, 107(45), pp. 19163-19166) and Weitz et al. (U.S. Patent Application Publication No. 2012 / 0132288). Generally, pico-injection is performed by injecting a seed into the capsule precursor, such as an emulsion droplet, before the capsule shell hardens, for example, before the capsule shell is formed.
[0129] The capsules may be of uniform or non-uniform size. In some cases, the diameter of the capsules may be approximately 0.001 μm, 0.01 μm, 0.05 μm, 0.1 μm, 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, 150 μm, 200 μm, 300 μm, 400 μm, 500 μm, 600 μm, 700 μm, 800 μm, 900 μm, or 1 mm. The capsules may have a diameter of at least about 0.001 μm, 0.01 μm, 0.05 μm, 0.1 μm, 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, 150 μm, 200 μm, 300 μm, 400 μm, 500 μm, 600 μm, 700 μm, 800 μm, 900 μm, or 1 mm. In some cases, the capsules may have a diameter of approximately 0.001 μm, 0.01 μm, 0.05 μm, 0.1 μm, 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, 150 μm, 200 μm, 300 μm, 400 μm, 500 μm, 600 μm, 700 μm, 800 μm, 900 μm, or less than 1 mm. In some cases, the capsules may have diameters of approximately 0.001 μm to 1 mm, 0.01 μm to 900 μm, 0.1 μm to 600 μm, 100 μm to 200 μm, 100 μm to 300 μm, 100 μm to 400 μm, 100 μm to 500 μm, 100 μm to 600 μm, 150 μm to 200 μm, 150 μm to 300 μm, or 150 μm to 400 μm.
[0130] The capsule may also have a specific density. In some cases, the capsule is less dense than an aqueous liquid (e.g., water); in some cases, the capsule is denser than an aqueous liquid. In some cases, the capsule is less dense than a non-aqueous liquid (e.g., oil); in some cases, the capsule is denser than a non-aqueous liquid. The capsule has a density of approximately 0.05 g / cm³. 3 , 0.1 g / cm³ 3 , 0.2 g / cm³ 3 0.3 g / cm³ 3 0.4 g / cm³ 3 , 0.5 g / cm 3 , 0.6 g / cm³ 3 , 0.7 g / cm³ 3 , 0.8 g / cm³ 3, 0.81 g / cm³ 3 , 0.82 g / cm³ 3 0.83 g / cm³ 3 0.84 g / cm³ 3 , 0.85 g / cm³ 3 , 0.86 g / cm³ 3 , 0.87 g / cm³ 3 0.88 g / cm³ 3 0.89 g / cm³ 3 0.90 g / cm³ 3 , 0.91 g / cm³ 3 , 0.92 g / cm³ 3 0.93 g / cm³ 3 0.94 g / cm³ 3 0.95 g / cm³ 3 , 0.96 g / cm³ 3 0.97 g / cm³ 3 , 0.98 g / cm³ 3 0.99 g / cm³ 3 , 1.00 g / cm³ 3 1.05 g / cm³ 3 , 1.1 g / cm³ 3 , 1.2 g / cm³ 3 1.3 g / cm³ 3 1.4 g / cm³ 3 1.5 g / cm³ 3 1.6 g / cm³ 3 1.7 g / cm³ 3 1.8 g / cm³ 3 1.9 g / cm³ 3 2.0 g / cm³ 3 , 2.1 g / cm³ 3 , 2.2 g / cm³ 3 2.3 g / cm³ 3 2.4 g / cm³ 3 , or 2.5 g / cm³ 3 The capsule may have a density of at least about 0.05 g / cm³. 3 , 0.1 g / cm³ 3 , 0.2 g / cm³ 3 , 0.3g / cm³ 3 , 0.4 g / cm³ 3 , 0.5 g / cm 3 , 0.6 g / cm³ 3 , 0.7 g / cm³ 3 , 0.8 g / cm³ 3 , 0.81 g / cm³3 , 0.82 g / cm³ 3 0.83 g / cm³ 3 0.84 g / cm³ 3 0.85 g / cm³ 3 , 0.86 g / cm³ 3 0.87 g / cm³ 3 0.88 g / cm³ 3 0.89 g / cm³ 3 0.90 g / cm³ 3 , 0.91 g / cm³ 3 , 0.92 g / cm³ 3 , 0.93 g / cm³ 3 0.94 g / cm³ 3 0.95 g / cm³ 3 , 0.96 g / cm³ 3 0.97 g / cm³ 3 , 0.98 g / cm³ 3 0.99 g / cm³ 3 , 1.00 g / cm³ 3 1.05 g / cm³ 3 , 1.1 g / cm³ 3 , 1.2 g / cm³ 3 1.3 g / cm³ 3 1.4 g / cm³ 3 1.5 g / cm³ 3 1.6 g / cm³ 3 1.7 g / cm³ 3 1.8 g / cm³ 3 1.9 g / cm³ 3 2.0 g / cm³ 3 , 2.1 g / cm³ 3 , 2.2 g / cm³ 3 2.3 g / cm³ 3 2.4 g / cm³ 3 , or 2.5 g / cm³ 3 It may have a density of . Otherwise, the capsule density is at most about 0.7 g / cm³. 3 , 0.8 g / cm³ 3 , 0.81 g / cm³ 3 , 0.82 g / cm³ 3 0.83 g / cm³ 3 0.84 g / cm³ 3 , 0.85 g / cm³ 3 , 0.86 g / cm³ 3 , 0.87 g / cm³ 30.88 g / cm³ 3 0.89 g / cm³ 3 0.90 g / cm³ 3 , 0.91 g / cm³ 3 , 0.92 g / cm³ 3 , 0.93 g / cm³ 3 0.94 g / cm³ 3 0.95 g / cm³ 3 , 0.96 g / cm³ 3 0.97 g / cm³ 3 , 0.98 g / cm³ 3 0.99 g / cm³ 3 , 1.00 g / cm³ 3 1.05 g / cm³ 3 , 1.1 g / cm³ 3 , 1.2 g / cm³ 3 1.3 g / cm³ 3 1.4 g / cm³ 3 1.5 g / cm³ 3 1.6 g / cm³ 3 1.7 g / cm³ 3 1.8 g / cm³ 3 1.9 g / cm³ 3 2.0 g / cm³ 3 , 2.1 g / cm³ 3 , 2.2 g / cm³ 3 2.3 g / cm³ 3 2.4 g / cm³ 3 , or 2.5 g / cm³ 3 This may also be the case. Such density may reflect the density of the capsule in any particular liquid (e.g., aqueous, water, oil, etc.).
[0131] 1. Capsule production by flow focusing In some cases, capsules can be produced by flow focusing. Flow focusing is a method of flowing a first liquid that does not mix with a second liquid into the second liquid. Referring to FIG. 12, a first (e.g., aqueous) liquid 1201 containing a monomer, a crosslinking agent, an initiator, and an aqueous surfactant is flowed into a second (e.g., oil) liquid 1202 containing a surfactant and an accelerator. After the second liquid enters the T-junction 1203 in the microfluidic device, the droplets of the first liquid break away from the flow of the first liquid, and due to the mixing of the monomer, crosslinking agent, and initiator in the first liquid with the accelerator in the second liquid, the capsule outer shell begins to form 1204. Thus, capsules are formed. As the capsules move downstream, the outer shell becomes thicker due to the increased exposure to the accelerator. Also, the thickness and permeability of the capsule outer shell can be varied using changes in reagent concentration.
[0132] Species, or other partitions such as droplets, can be encapsulated, for example, by including the species in the first liquid. Inclusion of the species in the second liquid can embed the species in the outer shell of the capsule. Of course, depending on the requirements of a particular sample processing method, the phases can be reversed, i.e., the first phase can be an oily phase and the second phase can be an aqueous phase.
[0133] 2. Production of Capsule-in-Capsule by Flow Focusing In some cases, the inner capsules can be produced by flow focusing. Referring to FIG. 13, a first (e.g., aqueous) liquid 1301 containing capsules, monomers, crosslinking agents, initiators, and an aqueous surfactant is flowed into a second (oil) liquid 1302 containing a surfactant and an accelerator. After the second liquid penetrates into the T-junction 1303 in the microfluidic device, the droplets of the first liquid break away from the flow of the first liquid, and due to the mixing of the monomers, crosslinking agents, and initiators in the first liquid with the accelerator in the second liquid, a second capsule outer shell begins to form around the capsule 1304. Thus, the inner capsule is formed. As the capsule progresses downstream, the outer shell becomes thicker due to the increased exposure to the accelerator. The thickness and permeability of the second capsule outer shell can also be varied using changes in reagent concentration.
[0134] The seeds can be encapsulated, for example, by including the seeds in the first liquid. Inclusion of the seeds in the second liquid can embed the seeds in the second outer shell of the capsule. Of course, depending on the requirements of a particular sample processing method, the phases can be reversed, i.e., the first phase can be an oil phase and the second phase can be an aqueous phase.
[0135] 3. Batch production of capsules In some cases, capsules can be produced batchwise using capsule precursors such as droplets in an emulsion. The capsule precursors can be formed by producing an emulsion by any suitable method, for example, using droplets containing monomers, crosslinking agents, initiators, and a surfactant. Then, an accelerator can be added to the medium to obtain capsule formation. With respect to the flow focusing method, the thickness of the outer shell can be varied by changing the concentration of the reactants and the exposure time to the accelerator. Then, the capsules can be washed and recovered. For any of the methods described herein, seeds containing other partitions can be encapsulated within the capsule or, where appropriate, within the outer shell.
[0136] In another example, droplets of the emulsion can be exposed to an accelerator present in the outlet well during the emulsion formation process. For example, capsule precursors can be formed by any preferred method, such as the flow focusing method illustrated in Figure 12. Rather than containing the accelerator in the second liquid 1202, the accelerator can be contained in the medium located at the outlet of the T-connection (e.g., the medium located at the right end of the horizontal channel in Figure 12). As the emulsion droplets (i.e., capsule precursors) exit the channel, they come into contact with the medium containing the accelerator (i.e., the outlet medium). If the capsule precursors have a density lower than that of the outlet medium, the capsule precursors rise through the medium, ensuring convective and diffusive exposure to the accelerator and reducing the likelihood of polymerization at the channel outlet.
[0137] VI. Species The methods, compositions, systems, apparatus, and kits of this disclosure can be used with any suitable species. The species may be any substance used in sample processing, such as reagents or analytes. Examples of species include whole cells, chromosomes, polynucleotides, organic molecules, proteins, polypeptides, carbohydrates, saccharides, sugars, lipids, enzymes, restriction enzymes, ligases, polymerases, barcodes, adapters, small molecules, antibodies, fluorophores, deoxynucleotide triphosphates (dNTPs), dideoxynucleotide triphosphates (ddNTPs), buffers, acidic solutions, basic solutions, temperature-sensitive enzymes, pH-sensitive enzymes, photosensitive enzymes, metals, metal ions, magnesium chloride, sodium chloride, manganese, aqueous buffers, mild buffers, ionic buffers, inhibitors, saccharides, oils, salts, ions, detergents, ionic detergents, nonionic detergents, oligonucleotides, nucleotides, DNA, RNA, peptide polynucleotides, complementary DNA (cDNA), and two Examples include stranded DNA (dsDNA), single-stranded DNA (ssDNA), plasmid DNA, cosmid DNA, chromosomal DNA, genomic DNA, viral DNA, bacterial DNA, mtDNA (mitochondrial DNA), mRNA, rRNA, tRNA, nRNA, siRNA, snRNA, snoRNA, scaRNA, microRNA, dsRNA, ribozymes, riboswitches and viral RNA, all or part locked nucleic acids (LNA), locked nucleic acid nucleotides, any other type of nucleic acid analog, proteases, nucleases, protease inhibitors, nuclease inhibitors, chelating agents, reducing agents, oxidizing agents, probes, chromophores, dyes, organic substances, emulsifiers, surfactants, stabilizers, polymers, water, small molecules, drugs, radioactive molecules, preservatives, antibiotics, and aptamers. In summary, the species used vary depending on the specific sample processing requirements.
[0138] In some cases, a partition contains a set of species having similar attributes (e.g., a set of enzymes, a set of minerals, a set of oligonucleotides, a mixture of different barcodes, a mixture of identical barcodes). In other cases, a partition contains a heterogeneous mixture of species. In some cases, a heterogeneous mixture of species contains all the components necessary to carry out a particular reaction. In some cases, such a mixture contains all the components necessary to carry out a reaction, except for one, two, three, four, five or more components necessary to carry out the reaction. In some cases, such additional components are contained within a different partition, or within or around the partition in a solution.
[0139] The species may be natural or synthetic. The species may be present in the sample obtained using any method known in the art. In some cases, the sample may be processed before it is analyzed for the analyte.
[0140] Seeds can be obtained from any suitable location, such as organisms, whole cells, cell preparations, and cell-free compositions derived from any organism, tissue, cell, or environment. Seeds can be obtained from environmental samples, biopsies, aspirates, formalin-fixed and embedded tissues, air, agricultural samples, soil samples, petroleum samples, water samples, or dust samples. In some examples, seeds can be obtained from bodily fluids, which may include blood, urine, feces, serum, lymph, saliva, mucosal secretions, sweat, central nervous system fluids, vaginal fluid, or semen. Seeds can also be obtained from manufactured products such as cosmetics, foods, and personal care products. Seeds may also be the product of experimental operations including recombinant cloning, polynucleotide amplification, polymerase chain reaction (PCR) amplification, purification methods (such as purification of genomic DNA or RNA), and synthesis reactions.
[0141] In some cases, the species may be quantified by mass. The species can be provided in masses of approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 50, 100, 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10000 ng, 1 μg, 5 μg, 10 μg, 15 μg, or 20 μg. The seeds can be provided in masses of at least approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 50, 100, 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10000 ng, 1 μg, 5 μg, 10 μg, 15 μg, or 20 μg. Seeds can be supplied in masses of approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 50, 100, 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10000 ng, 1 μg, 5 μg, 10 μg, 15 μg, or less than 20 μg. Seeds can be provided in masses ranging from approximately 1 to 10, 10 to 50, 50 to 100, 100 to 200, 200 to 1000, 1000 to 10000 ng, 1 to 5 μg, or 1 to 20 μg. If the seeds are polynucleotides, as described elsewhere in this disclosure, the amount of polynucleotides can be increased using amplification.
[0142] Furthermore, polynucleotides can also be quantified as "genome equivalents." A genome equivalent is the amount of polynucleotides equivalent to the haploid genome of the organism from which the target polynucleotide is induced. For example, a single diploid cell contains two genome equivalents of DNA. Polynucleotides can be supplied in amounts of genome equivalents ranging from approximately 1 to 10, 10 to 50, 50 to 100, 100 to 1000, 1000 to 10000, 10000 to 100000, or 100000 to 1000000. Polynucleotides are divided into at least approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 50, 100, 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500 It can be supplied in quantities of 9000, 9500, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, or 1000000 genome equivalents. Polynucleotides are divided into approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 50, 100, 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, and 9000. It can be supplied in quantities of 9500, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, or less than 1,000,000 of the genome equivalent.
[0143] Polynucleotides can also be quantified by the amount of sequence coverage provided. Sequence coverage refers to the average number of reads representing a given nucleotide in the reconstructed sequence. Generally, the more times a region is sequenced, the more accurate the sequence information obtained. Polynucleotides can be provided in amounts that provide a range of sequence coverage from approximately 0.1X to 10X, 10X to 50X, 50X to 100X, 100X to 200X, or 200X to 500X. Polynucleotides can be provided in amounts that provide sequence coverage of at least approximately 0.1X, 0.2X, 0.3X, 0.4X, 0.5X, 0.6X, 0.7X, 0.8X, 0.9X, 1.0X, 5X, 10X, 25X, 50X, 100X, 125X, 150X, 175X, or 200X. Polynucleotides can be provided in quantities that offer sequence coverage of approximately 0.2X, 0.3X, 0.4X, 0.5X, 0.6X, 0.7X, 0.8X, 0.9X, 1.0X, 5X, 10X, 25X, 50X, 100X, 125X, 150X, 175X, or less than 200X.
[0144] In some cases, the seeds are introduced into the partitions before or after a specific step. For example, a lysis buffer reagent can be introduced into the partitions after the partitioning of the cell sample into the partitions. In some cases, the reagents and / or partitions containing the reagents are introduced sequentially so that different reactions or operations occur in different steps. Alternatively, the reagents (or partitions containing the reagents) may be filled in steps where the reaction or operation steps are interspersed. For example, capsules containing a reagent for fragmenting molecules (e.g., nucleic acids) can be filled into wells, followed by a fragmentation step, and then capsules containing a reagent for ligating barcodes (or other unique identifiers, e.g., antibodies) can be filled, and then the barcodes can be ligated onto the fragmented molecules.
[0145] VII. Processing of Analytes and Other Species In some cases, the methods, compositions, systems, apparatus, and kits of this disclosure can be used to process a sample containing a species, such as an analyte. Any suitable process can be carried out.
[0146] a. Fragmentation of target polynucleotides In some cases, the methods, compositions, systems, apparatus, and kits of this disclosure can be used for polynucleotide fragmentation. Polynucleotide fragmentation is used as a step in various methods, including polynucleotide sequencing. Typically described in terms of length (quantified by the number of nucleotides per fragment), the size of the polynucleotide fragments may vary depending on the source of the target polynucleotide, the method used for fragmentation, and the desired application. A single fragmentation step or multiple fragmentation steps can be used.
[0147] The fragments produced using the method described herein may have a nucleotide length of about 1 to 10, 10 to 20, 20 to 50, 50 to 100, 50 to 200, 100 to 200, 200 to 300, 300 to 400, 400 to 500, 500 to 1000, 10000 to 50000, 5000 to 10000, 100000 to 100000, 100000 to 250000, or 250000 to 500000. The fragments produced using the method described herein may have a nucleotide length of at least about 10, 20, 100, 200, 300, 400, 500, 1000, 5000, 10000, 100000, 250000, or 500000 or more. The fragments produced using the method described herein may have nucleotide lengths of approximately 10, 20, 100, 200, 300, 400, 500, 1000, 5000, 10000, 100000, 250000, and less than 500000.
[0148] Fragments produced using the method described herein may have an average or median length of about 1 to 10, 10 to 20, 20 to 50, 50 to 100, 50 to 200, 100 to 200, 200 to 300, 300 to 400, 400 to 500, 500 to 1000, 10000 to 50000, 5000 to 10000, 100000 to 100000, 100000 to 250000, or 250000 to 500000 nucleotides. Fragments produced using the method described herein may have an average or median length of at least about 10, 20, 100, 200, 300, 400, 500, 1000, 5000, 100000, 100000, 250000, or 500000 or more nucleotides. The fragments produced using the method described herein may have an average or median nucleotide length of about 10, 20, 100, 200, 300, 400, 500, 1000, 5000, 10000, 100000, 250000, or less than 500000.
[0149] Several fragmentation methods are known in the art. For example, fragmentation can be carried out by physical, mechanical, or enzymatic methods. Physical fragmentation may include exposing the target polynucleotide to heat or UV light. By mechanical fracture, the target polynucleotide can be mechanically sheared into fragments of a desired range. Mechanical shearing can be achieved by several methods known in the art, including repeated pipetting of the target polynucleotide, sonication (e.g., using ultrasound), cavitation, and atomization. Alternatively, the target polynucleotide can be fragmented using enzymatic methods. In some cases, enzymatic digestion can be carried out using enzymes such as restriction enzymes.
[0150] The fragmentation methods described in the preceding paragraph and in some paragraphs of this disclosure are described with reference to a “target” polynucleotide, but this does not mean that they are limited above or elsewhere in this disclosure. Any fragmentation method described herein or known in the art can be applied to any polynucleotide used in conjunction with the present invention. In some cases, the polynucleotide may be a target polynucleotide, such as a genome. In other cases, the polynucleotide may be a fragment of a target polynucleotide that a person skilled in the art would like to further fragment. In yet other cases, further fragments may be further fragmented. Any suitable polynucleotide can be fragmented according to the methods described herein.
[0151] Restriction enzymes can be used to perform specific or nonspecific fragmentation of target polynucleotides. The methods of this disclosure may use one or more types of restriction enzymes, commonly described as type I enzymes, type II enzymes, and / or type III enzymes. Type II and type III enzymes are generally commercially available and well known in the art. Type II and type III enzymes recognize specific sequences of nucleotide base pairs ("recognition sequences" or "recognition sites") within a double-stranded polynucleotide sequence. Upon binding to and recognition of these sequences, type II and type III enzymes cleave the polynucleotide sequence. In some cases, the cleavage results in a polynucleotide fragment containing a protruding portion of single-stranded DNA called a "stick end." In other cases, the cleavage does not result in a fragment containing a protrusion, but produces a "blunt end." The methods of this disclosure may include the use of restriction enzymes that produce either a sticky end or a blunt end.
[0152] Restriction enzymes can recognize various recognition sites within a target polynucleotide. Some restriction enzymes ("precision cutters") recognize only a single recognition site (e.g., GAATTC). Other restriction enzymes are more heterogeneous and recognize more than one recognition site, or a variety of recognition sites. Some enzymes cleave at a single position within the recognition site, while others can cleave at multiple positions. Some enzymes cleave at the same position within the recognition site, while others cleave at variable positions.
[0153] The present disclosure provides a method of selecting one or more restriction enzymes to produce fragments of a desired length. Polynucleotide fragmentation is simulated in silico and optimized to obtain the greatest number or fraction of polynucleotide fragments within a particular size range while minimizing the number or fraction of fragments within an undesirable size range. An optimization algorithm can be applied to select a combination of two or more enzymes to yield a desired fragment size having a desired distribution of fragment amounts.
[0154] A polynucleotide can be exposed to two or more restriction enzymes simultaneously or sequentially. This can be accomplished, for example, by adding more than one restriction enzyme to a partition, or by adding one restriction enzyme to a partition, performing digestion, inactivating the restriction enzyme (e.g., by heat treatment), and then adding a second restriction enzyme. Any suitable restriction enzyme can be used alone or in combination in the methods presented herein.
[0155] In some cases, the species is a restriction enzyme that is a "rare-cutter." As used herein, the term "rare-cutter enzyme" generally refers to an enzyme having recognition sites that are only rarely present in the genome. The size of restriction fragments generated by cutting a hypothetical random genome with a restriction enzyme can be approximated by 4 N (where N is the number of nucleotides in the recognition site of the enzyme). For example, an enzyme having a recognition site consisting of 7 nucleotides has 4 7The genome is cut once per bp, producing a fragment of approximately 16,384 bp. Generally, rare cutter enzymes have a recognition site containing six or more nucleotides. For example, a rare cutter enzyme may contain or have a recognition site consisting of 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides. Examples of rare cutter enzymes include NotI (GCGGCCGC), XmaIII (CGGCCG), SstII (CCGCGG), SalI (GTCGAC), NruI (TCGCGA), NheI (GCTAGC), Nb.BbvCI (CCTCAGC), BbvCI (CCTCAGC), AscI (GGCGCGCC), AsiSI (GCGATCGC), FseI (GGCCGGCC), and PacI (TTAATTA). A), PmeI (GTTTAAAC), SbfI (CCTGCAGG), SgrAI (CRCCGGYG), SwaI (ATTTAAAT), BspQI (GCTCTTC), SapI (GCTCTTC), S fiI(GGCCNNNNNGGCC), CspCI(CAANNNNNGTGG), AbsI(CCTCGAGG), CciNI(GCGGCCGC), FspAI(RTGCGCAY), MauBI(CGC GCGCG), MreI (CGCCGGCG), MssI (GTTTAAAC), PalAI (GGCGCGCC), RgaI (GCGATCGC), RigI (GGCCGGCC), SdaI (CCTGCA GG), SfaAI (GCGATCGC), SgfI (GCGATCGC), SgrDI (CGTCGACG), SgsI (GGCGCGCC), SmiI (ATTTAAAT), SrfI (GCCCGGGC) Examples include Sse2321(CGCCGGCG), Sse83871(CCTGCAGG), LguI(GCTCTTC), PciSI(GCTCTTC), AarI(CACCTGC), AjuI(GAANNNNNNNTTGG), AloI(GAACNNNNNNTCC), BarI(GAAGNNNNNNTAC), PpiI(GAACNNNNNCTC), and PsrI(GAACNNNNNNTAC).
[0156] In some cases, polynucleotides can be simultaneously fragmented and barcoded. For example, polynucleotides can be fragmented using a transposase (e.g., NEXTERA), and a barcode can be attached to the polynucleotide.
[0157] VIII. Stimulus Responsiveness In some cases, stimuli can be used to induce the release of species from partitions. Generally, stimuli can cause disruption of the partition's structure, such as the well walls, spot components, droplet stability (e.g., droplets in emulsion), or capsule shell. These stimuli are particularly useful in inducing partitions to release their contents. Partitions may be contained within other partitions, and each partition may be responsive (or unresponsive) to different stimuli; therefore, stimulus responsiveness can be used to cause the contents of one partition (e.g., a partition that responds to a stimulus) to be released into another partition (e.g., a partition that does not respond to that stimulus or is less responsive to that stimulus).
[0158] In some cases, the contents of the inner capsule can be released into the contents of the outer capsule by applying a stimulus to dissolve the inner capsule, resulting in a capsule containing the mixed sample. Of course, this embodiment is purely illustrative, and the contents of any suitable partition can be released into any other suitable partition, medium, or container using stimulus responsiveness (see, for example, Table 1 for more specific examples of partitions within partitions).
[0159] Examples of stimuli that can be used include chemical stimuli, changes in volume, biological stimuli, light, temperature stimuli, magnetic stimuli, addition of media to the wells, and any combination thereof, as fully described below (see, for example, Esser-Kahn et al. (2011) Macromolecules 44: pp. 5539–5553; Wang et al. (2009) ChemPhysChem 10: pp. 2405–2409).
[0160] a. Chemical irritation and bulk change Partition breakdown can be induced using several chemical triggers (e.g., Plunkett et al., Biomacromolecules, 2005, 6: pp. 632–637). Examples of these chemical changes, but not limited to, include pH-mediated changes to the integrity of partition components, breakdown of partition components due to chemical cleavage of crosslinking bonds, and induced depolymerization of partition components. Partition breakdown can also be induced using bulk changes.
[0161] Changes in the pH of a solution, such as a decrease in pH, can induce partition breakdown through several different mechanisms. The addition of acid can cause partial decomposition or deassembly of partitions through various mechanisms. The addition of protons can deassemble polymer crosslinks in the partition components, break ionic or hydrogen bonds in the partition components, or create nanopores in the partition components, allowing the internal contents to leak out. Changes in pH can also destabilize emulsions, leading to the release of droplet contents.
[0162] In some cases, partitions are produced from materials containing acid-degradable chemical crosslinking agents such as ketals. A decrease in pH, particularly to a pH below 5, can induce the conversion of ketals to ketones and two alcohols, facilitating the breakdown of the partitions. In other cases, partitions can be produced from materials containing one or more pH-sensitive polyelectrolytes. A decrease in pH can break the ionic or hydrogen bonding interactions of such partitions or create nanopores within them. In some cases, partitions made from polyelectrolyte-containing materials include a core based on a charged gel that expands and contracts in response to changes in pH.
[0163] The fracture of crosslinked materials containing partitions can be achieved by several mechanisms. In some cases, the partitions can be brought into contact with various chemicals that induce oxidation, reduction, or other chemical changes. In some cases, reducing agents such as beta-mercaptoethanol can be used so that the disulfide bonds of the partitions are broken. Furthermore, the loss of partition integrity can be brought about by adding enzymes to cleave peptide bonds in the material forming the partitions.
[0164] Depolymerization can also be used to break up the partitions. Chemical triggers can be added to facilitate the removal of protective head groups. For example, the trigger can cause the removal of carbonate ester or carbamate head groups within the polymer, which then causes depolymerization and the release of seeds from inside the partitions.
[0165] In yet another example, the chemical trigger may include an osmotic trigger, where a change in the concentration of ions or solutes in the solution induces expansion of the material used to create the partition. This expansion can cause an increase in internal pressure, such that the partition bursts and releases its contents. The expansion can also cause an increase in the pore size of the material, diffusing the seeds contained within the partition, and vice versa.
[0166] Partitions can also be constructed to release their contents through changes in volume or physical properties, such as pressure-induced rupture, melting, or changes in porosity.
[0167] b. Biological stimulation Biological stimuli can also be used to induce partition disruption. Generally, biological triggers are similar to chemical triggers, but many examples use biomolecules or molecules commonly found in biological systems, such as enzymes, peptides, saccharides, fatty acids, and nucleic acids. For example, partitions can be made from materials containing polymers with peptide crosslinks that are sensitive to cleavage by specific proteases. More specifically, one example may involve partitions made from materials containing GFLGK peptide crosslinks. Upon addition of a biological trigger, such as the protease cathepsin B, the peptide crosslinks in the outer shell wells are cleaved, releasing the contents of the capsule. In other cases, the protease can be thermally activated. In another example, the partition contains a cellulose-containing component. The addition of the hydrolytic enzyme chitosan serves as a biological trigger for cleavage of cellulose bonds, depolymerization of the chitosan-containing component of the partition, and release of its internal contents.
[0168] c. Temperature stimulation Partitions can also be induced to release their contents upon application of a thermal stimulus. Temperature changes can cause a variety of changes in the partition. Thermal changes can cause melting of the partition or breakdown of the emulsion, such that part of the partition collapses. In other cases, heat can increase the internal pressure of the internal components of the partition, causing it to rupture or explode. In yet other cases, heat can deform the partition into a shrunk, dehydrated state. Heat can also act on heat-sensitive polymers used as materials to construct the partition.
[0169] In one example, the partition is made from a material containing a temperature-sensitive hydrogel. When heat is applied, such as at temperatures above 35°C, the hydrogel material shrinks. This sudden shrinkage of the material increases the pressure, causing the partition to burst.
[0170] In some cases, the material used to produce the partition may include diblock polymers or mixtures of two polymers having different heat sensitivities. One polymer may shrink after the application of heat, while the other may be more thermally stable. When heat is applied to such an outer shell wall, the heat-sensitive polymer will shrink, while the other will remain intact, which may cause the formation of pores. In yet other cases, the material used to produce the partition may include magnetic nanoparticles. Exposure to a magnetic field may cause heat generation, which may lead to the rupture of the partition.
[0171] d. Magnetic stimulation The inclusion of magnetic nanoparticles in the material used to produce partitions may enable the induced rupture of the partitions, as well as the induction of these partitions into other partitions (e.g., into wells in an array of capsules). In one example, the inclusion of Fe3O4 nanoparticles in the material used to produce partitions induces rupture in the presence of an oscillating magnetic field stimulus.
[0172] e. Electrical and optical stimulation Partitions can also be destroyed as a result of electrical stimulation. Similar to the magnetic particles described in the preceding section, electrosensitive particles can enable both induced partition rupture and other functions such as alignment in an electric field or redox reaction. In one example, a partition made from a material containing an electrosensitive material is aligned in an electric field such that the release of internal reagents can be controlled. In another example, an electric field can induce a redox reaction within the partition, which can increase its porosity.
[0173] Partitions can also be destroyed using light stimulation. Several photo-triggers are possible and may include systems using various molecules such as nanoparticles and chromophores that can absorb photons of a specific range of wavelengths. For example, a particular partition can be produced using a metal oxide coating. UV irradiation of a partition coated with SiO2 / TiO2 can result in the collapse of the partition wall. In yet another example, photo-switchable materials such as azobenzene groups can be incorporated into the material used to produce the partition. When UV or visible light is applied, these and other chemicals undergo reversible cis-trans isomerization upon photon absorption. In this embodiment, the incorporation of a photoswitch results in the collapse of part of the partition or an increase in the porosity of part of the partition.
[0174] f. Application of stimulation The apparatus, methods, compositions, systems, and kits of this disclosure can be used in combination with any instrument or device that provides such a trigger or stimulus. For example, if the stimulus is a thermal stimulus, the apparatus can be used in combination with a heat or temperature control plate that allows heating of the wells to induce capsule rupture. Several methods of heat transfer, such as the application of heat by radiative heat transfer, convective heat transfer, or conductive heat transfer, can be used for the thermal stimulus, but are not limited to these. In other cases, if the stimulus is a biological enzyme, the enzyme can be injected into the apparatus so that the enzyme is placed in each well. In another embodiment, if the stimulus is a magnetic or electric field, the apparatus can be used in combination with a magnetic or electric plate.
[0175] IX.Applications a. Determination of polynucleotide sequences In general, the methods and compositions provided herein are useful for the preparation of polynucleotide fragments for downstream applications such as sequencing. Sequencing can be performed by any preferred technique. For example, sequencing can be performed by the classical Sanger sequencing method. Other sequencing methods include high-efficiency sequencing, pyrosequencing, synthetic sequencing, single-molecule sequencing, nanopore sequencing, ligation sequencing, hybridization sequencing, RNA-Seq (Illumina), Digital Gene Expression (Helicos), next-generation sequencing, synthetic single-molecule sequencing (SMSS) (Helicos), large-scale parallel sequencing, clonal single-molecule arrays (Solexa), shotgun sequencing, Maxim-Gilbert sequencing, primer walking, and any other sequencing methods known in the industry.
[0176] In some cases, a varying number of fragments are sequenced. For example, in some cases, approximately 30% to 90% of the fragments are sequenced. In some cases, approximately 35% to 85%, 40% to 80%, 45% to 75%, 50% to 70%, 55% to 65%, or 50% to 60% of the fragments are sequenced. In some cases, at least approximately 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the fragments are sequenced. In some cases, less than approximately 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the fragments are sequenced.
[0177] In some cases, sequences derived from fragments are aggregated to provide sequence information about a contiguous region of the original target polynucleotide that is longer than the individual sequence reads. The individual sequence reads may have nucleotide lengths of approximately 10–50, 50–100, 100–200, 200–300, or 300–400 or more.
[0178] The identity of barcode tags can be useful for ordering sequence reads derived from individual fragments and for distinguishing between haplotypes. For example, during the compartmentalization of individual fragments, parent polynucleotide fragments can be separated into different partitions. As the number of partitions increases, the probability of fragments within the same partition originating from both the maternal and parent haplotypes becomes negligibly small. Therefore, sequence reads derived from fragments within the same partition can be aggregated and ordered.
[0179] b. Polynucleotide phase transformation This disclosure also provides methods and compositions for preparing polynucleotide fragments in a manner that can generate phase-shifted or ligation information. Such information can enable the detection of relevant genetic alterations in a sequence, such as genetic alterations separated by long stretches of polynucleotides (e.g., SNPs, mutations, indels, copy number changes, transversions, translocations, inversions, etc.). The term “indel” refers to co-localized insertions and deletions, as well as mutations resulting in a net increase or decrease of nucleotides. “Microindel” is an indel resulting in a net increase or decrease of 1 to 50 nucleotides. These alterations may exist in cis or trans relationships. In a cis relationship, two or more genetic alterations are present in the same polynucleotide or chain. In a trans relationship, two or more genetic alterations are present on multiple polynucleotide molecules or chains.
[0180] The polynucleotide phaseization can be determined using the methods provided herein. For example, a polynucleotide sample (e.g., a polynucleotide spread across a given locus or multiple loci) can be partitioned so that at most one polynucleotide molecule is present per partition. The polynucleotides can then be fragmented, barcoded, and sequenced. The sequences can be examined for genetic alterations. Detection of genetic alterations in the same sequence tagged with two different barcodes may indicate that the two genetic alterations are induced from two different strands of DNA and reflect a trans relationship. Conversely, detection of two different genetic alterations tagged with the same barcode may indicate that the two genetic alterations originate from the same strand of DNA and reflect a cis relationship.
[0181] Phase information can be important for characterizing polynucleotide fragments, particularly when the fragments originate from subjects who have or are suspected of having a specific disease or disorder (e.g., recessive genetic diseases such as cystic fibrosis, cancer, etc.). This information can distinguish between the following possibilities: (1) two genetic alterations within the same gene on the same strand of DNA and (2) two genetic alterations within the same gene but located on different strands of DNA. Possibility (1) may indicate that one copy of the gene is normal and the individual is disease-free, while possibility (2) may indicate that the individual has or will develop the disease, especially if the two genetic alterations impair gene function when they are located within the same gene copy. Similarly, phase information can also distinguish between the following possibilities: (1) two genetic alterations within different genes on the same strand of DNA and (2) two genetic alterations within different genes but located on different strands of DNA.
[0182] c. Sequencing of polynucleotides derived from a small number of cells The methods provided herein can also be used to prepare intracellular polynucleotides in a manner that allows for the acquisition of cell-specific information. This method enables the detection of genetic alterations (e.g., SNPs, mutations, indels, copy number changes, transversions, translocations, inversions, etc.) from very small samples, such as from samples containing about 10 to 100 cells. In some cases, at least about 1, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 cells can be used in the methods provided herein. In other cases, at most about 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 cells can be used in the methods provided herein.
[0183] In one example, the method includes partitioning a cell sample (or crude cell extract) such that at most one cell (or cell extract) is present per partition, lysing the cells, fragmenting the polynucleotides contained within the cells by any method described herein, attaching the fragmented polynucleotides to barcodes, pooling them, and sequencing them.
[0184] As described elsewhere in this specification, barcodes and other reagents can be contained within partitions (e.g., capsules). These capsules can be packed into separate partitions (e.g., wells) before, after, or concurrently with cell packing, so that each cell comes into contact with a different capsule. Using this technique, unique barcodes can be attached to polynucleotides obtained from each cell. The resulting tagged polynucleotides can then be pooled, sequenced, and their origins tracked using the barcodes. For example, polynucleotides with the same barcode can be determined to originate from the same cell, while polynucleotides with different barcodes can be determined to originate from different cells.
[0185] The methods described herein can be used to detect the distribution of cancer mutations across a population of cancer tumor cells. For example, some tumor cells may have mutations or amplifications of oncogenes (e.g., HER2, BRAF, EGFR, KRAS) in both alleles (homozygous), others may have mutations in one allele (heterozygous), and still others may not have mutations at all (wild-type). The methods described herein can be used to detect these differences and to quantify the relative numbers of homozygous, heterozygous, and wild-type cells. Such information can be used, for example, to grade specific cancers and / or to monitor cancer progression and its treatment over time.
[0186] In some cases, this disclosure provides a method for identifying mutations in two different oncogenes (e.g., KRAS and EGFR). If the same cell contains genes with both mutations, this may indicate a more aggressive form of cancer. In contrast, if the mutations are located in two different cells, this may indicate a more benign or less advanced form of cancer.
[0187] d. Analysis of gene expression The methods of this disclosure may be applicable to the processing of samples for the detection of changes in gene expression. The sample may include cells, mRNA, or cDNA reverse-transcribed from mRNA. The sample may be a pooled sample containing extracts from several different cells or tissues, or a sample containing an extract from a single cell or tissue.
[0188] Cells can be directly placed in a partition (e.g., a microwell) and lysed. After lysis, the polynucleotides of the cells can be fragmented and barcoded for sequencing using the method of the present invention. After extracting the polynucleotides from the cells, they can also be introduced into the partition used in the method of the present invention. Reverse transcription of mRNA can be carried out in or outside the partition described herein. CDNA sequencing can provide indication of the abundance of a particular transcript in a particular cell over time or after exposure to specific conditions.
[0189] The method presented above offers several advantages over current polynucleotide processing methods. Firstly, interoperator variability is significantly reduced. Secondly, the method can be performed in a low-cost, easily assembled microfluidic apparatus. Thirdly, controlled fragmentation of the target polynucleotide allows the user to produce polynucleotide fragments of a specified appropriate length. This helps in the compartmentalization of polynucleotides and also reduces the amount of sequence information loss due to the presence of excessively large fragments. The method and system also provide a straightforward workflow that maintains the integrity of the processed polynucleotide. Furthermore, the use of restriction enzymes allows the user to create DNA overhangs ("adhesion ends") that can be designed for compatibility with adapters and / or barcodes.
[0190] e. Separation of polynucleotides, such as chromosomes, from cells. In one example, polynucleotides, such as entire chromosomes, can be partitioned from cells using the methods, compositions, systems, apparatus, and kits provided herein. In one example, a single cell or a group of cells (e.g., 2, 10, 50, 100, 1000, 10000, 25000, 50000, 100000, 500000, 1,00000 or more cells) are packed into a container containing lysis buffer and proteinase K and incubated for a specified period. The use of multiple cells enables polynucleotide phase-ization, for example, by partitioning each polynucleotide and analyzing it in its own partition.
[0191] After incubation, the cell lysates are partitioned, for example, by flow focusing into capsules. If phase focusing is to be performed, flow focusing is carried out so that each capsule contains only a single analyte (e.g., a single chromosome) or only a single copy of any specific chromosome (e.g., one copy of chromosome 1 and one copy of chromosome 2). In some cases, multiple chromosomes can be encapsulated in the same capsule, as long as the chromosomes are not the same chromosome. Encapsulation is performed under gentle flow to minimize shearing of the polynucleotides. The capsules may be porous to allow washing of the capsule contents and introduction of reagents into the capsule while maintaining the polynucleotides (e.g., chromosomes) within the capsule. The encapsulated polynucleotides (e.g., chromosomes) can then be processed according to any method provided in this disclosure or known in the art. The capsule shell protects the encapsulated polynucleotides (e.g., chromosomes) from shearing and further degradation. Of course, this method can also be applied to any other cellular component.
[0192] As described above, the capsule shell can be used to protect polynucleotides from shearing. However, the capsule can also be used as a partition to allow compartmentalized shearing of the polynucleotide or other analytes. For example, in some cases, the polynucleotide can be encapsulated and then subjected to ultrasonic shearing or any other suitable shearing method. The capsule shell can be configured to remain intact under shearing, while the encapsulated polynucleotide can be sheared but retained within the capsule. In some cases, the same objective can be achieved using hydrogel droplets.
[0193] f. Detection of cancer mutations and forensic medicine The barcoding method using the partition-based amplification barcoding scheme described herein may be useful for generating barcode libraries from degraded samples, such as formalin-fixed paraffin-embedded (FFPE) tissue sections. The method described herein can identify that all amplicons within a partition originate from the same initial molecule. In fact, partition barcoding can be used to retain information about unique starting polynucleotides. Such identification can help determine the complexity of a library, as it can distinguish amplicons originating from different molecules. Furthermore, the method described herein allows for an assessment of unique inclusion and can help determine variant calling susceptibility. These advantages may be particularly useful in the detection of cancer mutations and in forensic science.
[0194] g. Low-input DNA application (circulating tumor cell (CTC) sequencing) The barcoding methods described herein may be useful in applications involving low polynucleotide inputs, such as sequencing of nucleic acids in circulating tumor cells (CTCs). For example, the MALBAC method described herein within a partition may help obtain good data quality in applications involving low polynucleotide inputs and / or filter out amplification errors.
[0195] X. Kit In some cases, the present disclosure provides a kit comprising reagents for generating partitions. The kit may include any suitable reagents as well as instructions for generating partitions and partitions within partitions.
[0196] In one example, the kit includes reagents for generating capsules within droplets in an emulsion. For example, the kit may include reagents for generating capsules, reagents for generating an emulsion, and instructions for introducing capsules into droplets of emulsion. Any suitable species can be incorporated into the droplets and / or capsules, as specified throughout this disclosure. The kits of this disclosure may also provide any of these species, such as polynucleotides with pre-differentiated barcodes. Similarly, as described throughout this disclosure, capsules can be designed to release their contents into droplets of emulsion upon application of a stimulus.
[0197] In another example, the kit includes reagents for generating an inner capsule. For example, the kit may include reagents for generating an inner capsule, reagents for generating an outer capsule, and instructions for generating an inner capsule. Any suitable species can be incorporated into the inner and / or outer capsules, as specified throughout this disclosure. The kits of this disclosure may also provide any of these species, such as polynucleotides with pre-segmented barcodes. Similarly, as described throughout this disclosure, the inner capsule can be designed to release its contents into the outer capsule upon application of a stimulus.
[0198] XI. Equipment In some cases, the present disclosure provides an apparatus including partitions for processing analytes. The apparatus may be a microwell array or a microspot array, as described elsewhere in the present disclosure. The apparatus can be formed in a manner that includes any preferred partitions. In some cases, the apparatus includes multiple wells or multiple spots. Of course, any partition in the apparatus may also hold other partitions, such as capsules or droplets in emulsions.
[0199] The apparatus can be formed from any suitable material. In some examples, the apparatus is formed from a material selected from the group consisting of quartz glass, soda-lime glass, borosilicate glass, poly(methyl methacrylate), sapphire, silicon, germanium, cyclic olefin copolymer, polyethylene, polypropylene, polyacrylate, polycarbonate, plastic, and combinations thereof.
[0200] In some cases, the apparatus includes channels for the flow of liquid into and between partitions. Any suitable channels can be used. The apparatus may also include liquid inlets and liquid outlets. The inlets and outlets can be coupled to a liquid handling device to introduce seeds into the apparatus. The apparatus can be sealed before or after the introduction of any seeds.
[0201] Hydrophilic and / or hydrophobic materials can be used in different parts of the apparatus. For example, in some cases the apparatus of the present disclosure includes a partition having an internal surface containing a hydrophilic material. In some cases the surface that is external to the partition contains a hydrophobic material. In some cases the liquid flow path is coated with a hydrophobic or hydrophilic material.
[0202] As will be understood, this disclosure provides the use of any composition, library, method, apparatus, and kit described herein for a specific use or purpose, including the various uses, applications, and purposes described herein. For example, this disclosure provides the use of the composition, method, library, apparatus, and kit described herein in species segmentation, oligonucleotide segmentation, stimulus-selective release of species from partitions, carrying out reactions in partitions (e.g., ligation and amplification reactions), carrying out nucleic acid synthesis reactions, nucleic acid barcoding, preparation of polynucleotides for sequencing, polynucleotide sequencing, polynucleotide phase formation, sequencing of polynucleotides from a small number of cells, gene expression analysis, polynucleotide segmentation from cells, mutation detection, neurological disorder diagnosis, diabetes diagnosis, fetal aneuploidy diagnosis, cancer mutation detection and forensic medicine, disease detection, medical diagnosis, low-input nucleic acid applications such as circulating tumor cell (CTC) sequencing, in combination thereof, and any other uses, methods, processes, or uses described herein. [Examples]
[0203] (Example 1) Production of fork-type adapter libraries containing barcode sequences by asymmetric PCR and addition of partially complementary universal sequences. This embodiment provides a method for manufacturing a fork-type adapter containing a barcode sequence compatible with next-generation sequencing technology (e.g., ILLUMINA). In this embodiment, the barcode is positioned at position 207, as shown in Figure 2.
[0204] Referring to Figure 4, a single-stranded adapter-barcode polynucleotide sequence 401 is synthesized, comprising a first immobilization region 402, a barcode region 403, and a first sequencing primer region 404. The barcode region 403 is a random sequence of 7 nucleotides synthesized by incorporating equimolar concentrations of A, G, T, and C during each coupling step.
[0205] After synthesis, the single-stranded adapter-barcoded polynucleotide 401 is diluted in aqueous droplets in an oil-in-water emulsion so that each droplet contains an average of 0.1 polynucleotides. The droplets also contain reagents (e.g., polymerase, primers, dNTPs, buffers, salts) and DNA insertion dyes (e.g., ethidium bromide) for amplification of the single-stranded adapter-barcoded polynucleotide 401 by asymmetric PCR. The reverse primer is present in excess of the forward primer, or vice versa, enabling asymmetric amplification. The polynucleotide is amplified, and the reaction proceeds through the exponential phase of amplification 410 to produce the double-stranded product 405, and then through the linear phase of amplification 411 to produce the single-stranded product 406.
[0206] The droplets are sorted on a fluorescence-assisted cell sorting system (FACS) 412 for collecting droplets containing amplified polynucleotides. A partially complementary universal sequence 407 is added to the partition to generate a partially annealed fork-type structure 413. The partially complementary universal sequence 407 includes a second immobilization region 408 and a second sequencing primer region 409, the latter including a T protrusion that fits with an A protrusion on the polynucleotide target to be sequenced (not shown).
[0207] (Example 2) Fragmentation and barcoding using fragmentation enzymes A single-stranded adapter-barcode polynucleotide sequence (e.g., 401 in Figure 4) comprising a first immobilization region 402, a barcode region 403, and a first sequencing primer region 404 is synthesized, segmented, amplified, and sorted as described in Example 1 or any other method described herein. Interfacial polymerization is carried out on droplets containing the single-stranded adapter-barcode polynucleotide sequence to produce a plurality of capsules containing a library of single-stranded adapter-barcode polynucleotide sequences 406, where each (or many) sequence in the library differs in the sequence of its corresponding barcode region 403. Thus, a library of encapsulated single-stranded adapter-barcode polynucleotides is produced.
[0208] Two mixtures are prepared. Mixture Z1 comprises a target polynucleotide (i.e., the polynucleotide to be fragmented and barcoded), a fragmentation enzyme (e.g., NEBNEXT DSDNA FRAGMENTASE), and a partially complementary universal sequence (e.g., 407 in Figure 4). The second mixture Z2 comprises a library of encapsulated single-stranded adapter-barcoded polynucleotides produced as described above, and magnesium chloride at a concentration sufficient to activate the fragmentation enzyme. Mixtures Z1, Z2, or both Z1 and Z2 also contain T4 polymerase, Taq polymerase, and a heat-stable ligase.
[0209] Mixtures Z1 and Z2 are mixed to form an intracapsule according to methods described elsewhere in this disclosure, such as flow focusing. Figure 5 illustrates an intracapsule produced according to the above method. The outer capsule 501 comprises the inner capsule 502 and the medium 504. The inner capsule 502 is one member of a library of encapsulated single-stranded adapter-barcode polynucleotides. Thus, the inner capsule 502 contains multiple copies of the single-stranded adapter-barcode polynucleotide 503, which can be used to encapsulate the same barcode on polynucleotides in partitions such as the outer capsule 501.
[0210] Medium 504 contains the contents of the above mixtures Z1 and Z2. More specifically, Medium 504 contains a target polynucleotide 505, a partially complementary universal sequence 506, and an enzyme mix 507 comprising a fragmentase, T4 polymerase, Taq polymerase, a heat-stable ligase, magnesium chloride, and a suitable buffer.
[0211] During the formation of the capsule-in-capsule and exposure of the capsule-in-capsule to appropriate conditions, enzymes process the target polynucleotide. More specifically, fragmentase fragments the target polynucleotide, and T4 polymerase blunts the ends of the fragmented target polynucleotide. The fragmentase and T4 polymerase are then thermally inactivated, and the inner capsule 502 is ruptured using a stimulus, releasing its contents into the outer capsule 501. Taq polymerase adds 3'-A protrusions to the fragmented and blunt-terminated target polynucleotide. The single-stranded adapter-barcoded polynucleotide 503 hybridizes with the partially complementary universal sequence 506 to form a fork-shaped adapter having a 3'-T protrusion that fits with the 3'-A protrusion on the fragmented target polynucleotide. A thermally stable ligase ligates the fork-shaped adapter to the fragmented target polynucleotide, producing a barcoded target polynucleotide. The outer capsule 501 is then ruptured, and the sample from all the outer capsules is pooled and the target polynucleotide is sequenced. Subsequently, prior to sequencing, additional preparation steps (e.g., bulk amplification, size selection, etc.) may be performed as needed.
[0212] In some cases, mixture Z1 contains multiple versions of a partially complementary universal sequence 506, where each version has its own sample-specific barcode.
[0213] Furthermore, while the above example uses a thermostable ligase to bind a fork-shaped adapter containing a barcode sequence to a target polynucleotide, this step can also be achieved using PCR, as described elsewhere in this disclosure.
[0214] (Example 3) Fragmentation and barcoding by ultrasonic processing A library of encapsulated single-stranded adapter-barcode polynucleotides is prepared as described in Example 2, or by any other preferred method described in this disclosure. The target polynucleotide (i.e., the polynucleotide to be fragmented) is fragmented into capsules. The capsules containing the target polynucleotide are configured to withstand ultrasonic stress. The capsules containing the target polynucleotide are exposed to ultrasonic stress (e.g., a COVARIS Focused-Ultrasonicator) to fragment the target polynucleotide and produce fragmented target polynucleotide capsules.
[0215] A mixture Z1 is prepared comprising a library of encapsulated single-stranded adapter-barcoded polynucleotides (e.g., Figure 4:406), fragmented target polynucleotide capsules, a partially complementary universal sequence (e.g., Figure 4:407), an enzyme mixture (T4 polymerase, Taq polymerase, and a heat-stable ligase), and a suitable buffer. Capsule-in-capsules are generated according to methods described elsewhere in this disclosure, such as flow focusing.
[0216] Figure 6 illustrates an example of an inner capsule produced according to the method described above. The outer capsule 601 comprises a plurality of inner capsules 602 and 605 and a medium 604. The inner capsules 602 and 605 each contain a capsule containing a single-stranded adapter-barcode polynucleotide 603 and a capsule containing a fragmented target polynucleotide 606, respectively. The inner capsule 602 contains multiple copies of the single-stranded adapter-barcode polynucleotide 603, which can be used to ligate the same barcode to polynucleotides within a partition, such as the fragmented target polynucleotide 606 contained in the inner capsule 605.
[0217] Medium 604 contains the contents of the above mixture Z1. More specifically, medium 604 contains a partially complementary universal sequence 607, an enzyme mixture (T4 polymerase, Taq polymerase, and a heat-stable ligase) 608, and a suitable buffer.
[0218] The inner capsule 605 containing the fragmented target polynucleotide 606 is exposed to a stimulus to rupture, releasing its contents into the contents of the outer capsule 601. T4 polymerase blunts the ends of the fragmented target polynucleotide; Taq polymerase adds 3'-A protrusions to the fragmented, blunt-ended target polynucleotide. The T4 polymerase and Taq polymerase are then thermally inactivated, and a stimulus is applied to release the contents of the inner capsule 602 into the outer capsule 601. The single-stranded adapter-barcoded polynucleotide 603 hybridizes with the partially complementary universal sequence 607 to form a fork-shaped adapter having a 3'-T protrusion that fits with the 3'-A protrusion on the fragmented target polynucleotide. A thermally stable ligase ligates the fork-shaped adapter to the fragmented target polynucleotide, producing a barcoded target polynucleotide. The outer capsule 601 is then ruptured, the sample from all outer capsules is pooled, and the target polynucleotide is sequenced.
[0219] As described in Example 2, in some cases Z1 may contain multiple versions of the partially complementary universal sequence 607. Furthermore, although this example demonstrates barcoding of the target polynucleotide by using a heat-stable ligase, this step can also be achieved using PCR.
[0220] (Example 4) Fork-type adapter generation by single-primer isothermal amplification (SPIA) and restricted digestion This embodiment demonstrates the synthesis of a fork-type adapter by SPIA and restriction digestion. Figure 7 provides an example of a product (or intermediate) that can be produced according to the method of this embodiment. Referring to Figure 7, a hairpin adapter 701 (SEQ ID NO: 2) is shown, which can be used as a precursor for a fork-type adapter as described elsewhere in this disclosure. In this embodiment, the hairpin adapter is synthesized as a single-strand amplification product using SPIA. The hairpin adapter 701 includes a double-stranded region 702, a 3'-T protrusion 703 for AT ligation, and a region 704 (i.e., between positions 33 and 34) that can be cleaved by restriction enzymes. The hairpin adapter may also include a barcode region and functional regions such as an immobilization region and a region for annealing of sequencing primers.
[0221] Cutting the adapter (e.g., between positions 33 and 34) generates a fork-shaped adapter as shown in Figure 8a (SEQ ID NOs: 3-4). This adapter is cut by introducing an oligonucleotide sequence complementary to the region to be cut and exposing the annealed adapter to restriction enzymes. Ligation of the fork-shaped adapter region shown in Figure 8a to the target polynucleotide results in the structure shown in Figure 8b (SEQ ID NOs: 5-6). Referring to Figure 8, the underlined portion of the sequence in Figure 8b contains a target polynucleotide having a 3'-A protrusion that fits with ligation to the fork-shaped adapter shown in Figure 8a.
[0222] Next, the sequences shown in Figure 8b (SEQ ID NOs. 5-6) are amplified by polymerase chain reaction to produce SEQ ID NOs. 7 (amplified product of SEQ ID NOs. 5) and SEQ ID NOs. 8 (amplified product of SEQ ID NOs. 6), as shown in Figure 8c. In Figure 8c, SEQ ID NOs. 7 is an amplified product of SEQ ID NOs. 5, with a first immobilized sequence (underlined 5' portion) and a second immobilized sequence (underlined 3' portion) added to SEQ ID NOs. SEQ ID NOs. 8 is an amplified product of SEQ ID NOs. 6, with the non-hybridized portion of SEQ ID NOs. 6 replaced by different sequences (underlined 3' portion and underlined 5' portion). Furthermore, SEQ ID NOs. 8 contains six nucleotide barcodes (TAGTGC; bold) within the 5' non-hybridized region of the polynucleotide. Therefore, the amplified product includes a barcoded target polynucleotide sequence (represented by 111), an immobilized sequence, and barcodes.
[0223] (Example 5) Further fork-type adapters using single-primer isothermal amplification (SPIA) and restricted digestion. This embodiment demonstrates the synthesis of the fork-shaped adapters (SEQ ID NOs: 9-10) shown in Figure 9a by SPIA and restricted digestion, where N represents A, T, G, or C. Figure 9b shows a single-strand fork-shaped adapter (SEQ ID NO: 11), where the single-strand form can create a hairpin structure. By cutting the hairpin structure at the positions indicated by the asterisks, the fork-shaped adapter shown in Figure 9a is obtained.
[0224] The template for SPIA is the sequence shown in Figure 9c (SEQ ID NO: 12). In Figure 9c, "R" represents an RNA region. Figure 9d shows the hairpin structure formed by the sequence in Figure 9c. The sequence in Figure 9d (SEQ ID NO: 12) is treated with polymerase to add a nucleotide to the 3' end, generating the sequence shown in Figure 9e (SEQ ID NO: 13). Next, the sequence in Figure 9e (SEQ ID NO: 13) is treated with RNase H, which degrades RNA hybridized to DNA, to obtain the sequence in Figure 9f (SEQ ID NO: 14).
[0225] Next, chain substitution SPIA is performed on sequence number 14. The primer for chain substitution amplification is RRRRRRRRRRRRR (i.e., R 13 This primer is in the form of ). This primer is an RNA primer that is one nucleotide longer than the non-hybridizing 3' end of SEQ ID NO: 14 (i.e., N 12 (Figure 9f). More specifically, as shown in Figure 9f, the 3' end of SEQ ID NO: 14 contains 12 N nucleotides. The RNA primer contains 13 nucleotides. Nucleotides 2-13 of the RNA primer are complementary to the 12 non-hybridized N nucleotides of SEQ ID NO: 14. Nucleotide 1 of the RNA primer is complementary to the first hybridized base (from 3' to 5'), in this case, T. The RNA primer substitutes A, producing the double-stranded extension product shown in Figure 9g (SEQ ID NOs: 15-16). Because only one primer is present, the reaction produces multiple copies of the single-stranded product. The single-stranded amplification product is treated with RNase H to produce the single-stranded amplification product shown in Figure 9h (SEQ ID NO: 17). Figure 9i shows this sequence (SEQ ID NO: 17) in 5'-3' form. Figure 9j shows this sequence (SEQ ID NO: 17) in hairpin form.
[0226] Next, the hairpin adapter shown in Figure 9j is ligated to a fragmented polynucleotide having a 3'-A protrusion. The hairpin is then cleaved between the A and C residues separated by the curve in Figure 9j by adding an oligonucleotide complementary to its region and cleaving with a restriction enzyme. This generates a fork-shaped adapter. Next, PCR amplification is performed as described in Example 4 to attach the immobilized region and barcode to the fork-shaped adapter bound to the target polynucleotide.
[0227] (Example 6) Generation of fork-type adapters containing barcodes by exponential PCR and hybridization. This embodiment demonstrates the production of a fork-type adapter containing a barcode by hybridization. Figure 10a shows an exemplary fork-type adapter provided in Figure 8a. As described in Example 4, this adapter can be ligated to a target polynucleotide, and then an amplification reaction can be carried out to add further functional sequences, including a barcode. However, it is also possible to directly incorporate the barcode (and other functional sequences) into the fork-type adapter and then bind the fork-type adapter to the target polynucleotide. For example, Figure 10b shows the fork-type adapter of Figure 10a with the addition of a first immobilization region (underlined) and a 7-nucleotide barcode region (bold / underlined; "N").
[0228] The barcoded fork-shaped adapter shown in Figure 10b is produced by first synthesizing Sequence ID No. 18 as a single strand. Diversity in the barcode region is generated using an equimolar mixture of A, G, T, and C, as described throughout this disclosure. Droplet-based PCR is performed as described in Example 1. However, Sequence ID No. 18 in the droplet is amplified using one DNA primer and one RNA primer. Amplification is performed in the presence of an insertion dye, and the droplet containing the amplified Sequence ID No. 18 is isolated as described in Example 1. Figure 10c shows the double-stranded amplification product. The underlined portion of Sequence ID No. 19 is the RNA strand derived from the RNA primer. The sequence shown in Figure 10c is then treated with RNase H, which digests the underlined RNA region, to obtain the construct shown in Figure 10d. To generate a fork-shaped construct, a partially complementary universal sequence (Sequence ID No. 21) is added to the construct shown in Figure 10d to produce the product shown in Figure 10e. The advantage of using this process is that it utilizes the significantly higher amplification of polynucleotides provided by exponential PCR compared to the linear amplification of polynucleotides provided by SPIA.
[0229] (Example 7) Double indexing method This embodiment demonstrates a method for synthesizing barcodes for dual-index reading. Dual-index reading involves reading both strands of a double-stranded fragment using barcodes linked to each strand. Figure 11 shows an example of barcode synthesis for the dual-index method and an example of using barcodes within a capsule in a capsule configuration.
[0230] As shown in Figure 11a, a first single-stranded adapter-barcode polynucleotide sequence 1101 is synthesized, comprising a first immobilization region 1102, a first barcode region 1103, and a first sequencing primer region 1104. In parallel, as shown in Figure 11b, a second single-stranded adapter-barcode polynucleotide sequence 1131 is synthesized, comprising a second immobilization region 1132, a second barcode region 1133, and a second sequencing primer region 1134. In some cases, barcode regions 1103 and 1133 are the same sequence. In other cases, barcode regions 1103 and 1133 are different sequences or partially different sequences.
[0231] After synthesis, single-stranded adapter-barcoded polynucleotides 1101 (Figure 11a) and 1131 (Figure 11b) are simultaneously diluted in aqueous droplets in an oil-in-water emulsion. The droplets also contain reagents (e.g., polymerase, primers, dNTPs, buffers, salts) and DNA insertion dyes (e.g., ethidium bromide) for amplification of single-stranded adapter-barcoded polynucleotides 1101 (Figure 11a) and 1131 (Figure 11b), respectively, by asymmetric PCR. Reverse primers are present in excess of forward primers, or vice versa, enabling asymmetric amplification. Polynucleotides 1101 (Figure 11a) and 1131 (Figure 11b) are amplified, and the reaction proceeds through the exponential phase of amplification 1110, producing double-stranded products 1105 (Figure 11a) and 1135 (Figure 11b), and then proceeds through the linear phase of amplification 1111, producing single-stranded products 1106 (Figure 11a) and 1136 (Figure 11b), respectively.
[0232] The droplets are sorted on a fluorescence-assisted cell sorting system (FACS) 1112 to collect the droplets containing amplified polynucleotides.
[0233] Next, interfacial polymerization is carried out on droplets containing droplets of single-stranded adapter-barcode polynucleotide sequences 1106 and 1136, respectively, to produce two types of capsules 1120 (Figure 11a) and 1150 (Figure 11b), each containing one of the single-stranded adapter-barcode polynucleotide sequences 1106 or 1136, respectively.
[0234] Two mixtures are prepared. Mixture Z1 contains the target polynucleotide (i.e., the polynucleotide to be fragmented and barcoded) 1170 and a fragmentation enzyme (e.g., NEBNEXT DSDNA FRAGMENTASE). The second mixture Z2 contains the capsules 1120 and 1180 produced as described above and magnesium chloride in a concentration sufficient to activate the fragmentation enzyme. Mixtures Z1, Z2, or both Z1 and Z2 also contain T4 polymerase, Taq polymerase, and a heat-stable ligase.
[0235] Mixtures Z1 and Z2 are mixed and allowed to form an intracapsule according to a method described elsewhere in this disclosure, such as flow focusing. Figure 11c illustrates an intracapsule produced according to the above method. The outer capsule 1160 comprises capsules 1120 and 1150 and a medium 1190. Thus, capsules 1120 and 1150 each contain multiple copies of single-stranded adapter-barcode polynucleotides 1106 and 1136, respectively, which can be used to conjugate barcodes 1103 and 1133 to polynucleotides in a partition, such as target polynucleotide 1170 in the medium 1190 of the outer capsule 1160.
[0236] Medium 1190 contains the contents of the above mixtures Z1 and Z2. More specifically, Medium 1190 contains the target polynucleotide 1170 and an enzyme mix 1180 comprising a fragmentase, T4 polymerase, Taq polymerase, a heat-stable ligase, magnesium chloride, and a suitable buffer.
[0237] During the formation of the capsule-in-capsule and exposure of the capsule-in-capsule to appropriate conditions, the enzymes process the target polynucleotide. More specifically, fragmentase fragments the target polynucleotide, and T4 polymerase blunts the ends of the fragmented target polynucleotide. The fragmentase and T4 polymerase are then thermally inactivated, and the capsules 1120 and 1150 are ruptured using stimulation, releasing their contents into the medium 1190 of the outer capsule 1160. Taq polymerase adds 3'-A protrusions to the fragmented, blunt-terminated target polynucleotide. Single-stranded adapter-barcode polynucleotide 1106 hybridizes with single-stranded adapter-barcode polynucleotide 1136 to form a fork-shaped adapter containing barcode regions 1103 and 1133, having 3'-T protrusions that fit with the 3'-A protrusions (not shown) on the fragmented target polynucleotide. A heat-stable ligase ligates the fork-shaped adapter to the fragmented target polynucleotide, generating a barcoded target polynucleotide. The outer capsule 1160 is then ruptured, and the sample from all outer capsules is pooled and the target polynucleotide is sequenced.
[0238] Furthermore, while the above example uses a thermostable ligase to bind a fork-shaped adapter containing a barcode sequence to a target polynucleotide, this step can also be achieved using PCR, as described elsewhere in this disclosure.
[0239] (Example 8) Production of fork-type adapters containing barcode sequences by bead emulsion PCR and addition of partially complementary universal sequences. As shown in Figure 14a, a single-stranded adapter-barcode sequence 1401 is synthesized, comprising a first immobilization region 1402, a barcode region 1403, and a first sequencing primer region 1404. After synthesis, the single-stranded adapter-barcode sequence 1401 is diluted in aqueous droplets in an oil-in-water emulsion so that each droplet contains an average of one polynucleotide. The droplets also contain a first bead 1405 linked by a photounstable linker to one or more copies of an RNA primer 1406 complementary to the sequence contained in the first sequencing primer region 1404; a DNA primer complementary to the sequence contained in the first immobilization region 1402 (not shown); and reagents necessary for amplification of the single-stranded adapter-barcode sequence 1401 (e.g., polymerase, dNTPs, buffer, salt). Both 1407 and 1407 bind to the first bead 1405 to form structure 1420, amplifying the polynucleotide in solution (not shown) to produce the double-stranded product 1408.
[0240] Next, the emulsion is broken up and the emulsion components are pooled to form a product mixture. As shown in Figure 14b, the freed beads are then washed several times in a suitable medium 1409 (by centrifugation), treated with sodium hydroxide (NaOH) 1410 to denature the double-stranded product bound to the first beads 1405, and then washed again 1411. After denaturation 1410 and washing 1411 of structure 1420, the resulting structure 1430 contains a single-stranded complement 1412 to the single-stranded adapter-barcode sequence 1401, which includes a complementary immobilization region 1413, a complementary barcode region 1414, and a complementary sequencing primer region 1415. As shown, the complementary sequencing primer region 1415 contains an RNA primer 1406. Structure 1430 is then resuspended in a suitable medium.
[0241] Next, as shown in Figure 14c, a second bead 1416 containing a complementary immobilization region 1413 and one or more copies of DNA polynucleotide 1417 complementary to it is added to the medium. The second bead 1416 binds to the single-stranded complement 1412 via the complementary immobilization region 1413 of the complementary DNA polynucleotide 1417 and the single-stranded complement 1412. Here, the single-stranded complement binds to one end of the first bead 1405, and the second bead 1416 forms structure 1440 at the other end.
[0242] As shown in Figure 14d, structure 1440 is then centrifuged 1418 using a glycerol gradient to separate structure 1440 from structure 1430 that is not contained in structure 1440. If the second bead 1416 is a magnetic bead, magnetic separation can be used as an alternative. The product is then treated with NaOH 1419 to denature the single-stranded complement 1412 derived from the second bead 1416 and obtain the regeneration of structure 1430. Structure 1430 is then subjected to several washes (by centrifugation) to remove the second bead 1416. The single-stranded complement 1412 bound to structure 1430 is a single-stranded barcode adapter.
[0243] As shown in Figure 14e, a fork-type adapter can be generated using single-stranded complement 1412. Next, to generate the fork-type adapter 1450, single-stranded complement 1412 is released from structure 1430 using light 1424 and then mixed with the universal complementary sequence 1426 1425, or mixed with the universal complementary sequence 1426 first 1425 and then released from structure 1430 1424. To generate ligable ends, the RNA primer 1406 of single-stranded complement 1412 is digested using RNAase H, and a single-base T overhang on the universal complementary sequence 1426 is generated using a type II restriction enzyme. The T overhang aligns with the A overhang on the polynucleotide target to be sequenced (not shown).
[0244] (Example 9) Production of fork-type adapters containing barcode sequences by bead emulsion PCR and addition of partially complementary universal sequences. As shown in Figure 15a, a single-stranded adapter-barcode sequence 1501 is synthesized, comprising a first immobilization region 1502, a barcode region 1503, and a first sequencing primer region 1504. After synthesis, the single-stranded adapter-barcode sequence 1501 is diluted in aqueous droplets in an oil-in-water emulsion so that each droplet contains an average of one polynucleotide. The droplets also contain a first bead 1505 linked by a photounstable linker to one or more copies of an RNA primer 1506 complementary to the sequence contained in the first immobilization region 1502; a DNA primer complementary to the sequence contained in the first sequencing primer region 1502 (not shown); and reagents necessary for amplification of the single-stranded adapter-barcode sequence 1501 (e.g., polymerase, dNTPs, buffers, salts). Both 1507 and 1507 bind to the first bead 1505 to form structure 1520, amplifying the polynucleotide in solution (not shown) to produce the double-stranded product 1508.
[0245] Next, the emulsion is broken up and the emulsion components are pooled to form a product mixture. As shown in Figure 15b, the freed beads are then washed several times in a suitable medium 1509 (by centrifugation), treated with sodium hydroxide (NaOH) 1510 to denature the double-stranded product bound to the first beads 1505, and then washed again 1511. After denaturation 1510 and washing 1511 of structure 1520, the resulting structure 1530 contains a single-stranded complement 1512 to the single-stranded adapter-barcode sequence 1501, which includes a complementary immobilization region 1513, a complementary barcode region 1514, and a complementary sequencing primer region 1515. As shown, the complementary sequencing primer region 1515 contains an RNA primer 1506. Structure 1530 is then resuspended in a suitable medium.
[0246] Next, as shown in Figure 15c, a second bead 1516 containing one or more copies of DNA polynucleotide 1517 complementary to the complementary sequencing primer region 1515 is added to the medium. The second bead 1516 binds to the single-stranded complement 1512 via the complementary sequencing primer region 1515 of the complementary DNA polynucleotide 1517 and the single-stranded complement 1512. Here, the single-stranded complement binds to one end of the first bead 1505, and the second bead 1516 forms structure 1540 at the other end.
[0247] As shown in Figure 15d, structure 1540 is then centrifuged 1518 using a glycerol gradient to separate structure 1540 from structure 1530 that is not contained in structure 1540. If the second bead 1516 is a magnetic bead, magnetic separation can be used as an alternative. The product is then treated with NaOH 1519 to denature the single-stranded complement 1512 derived from the second bead 1516 and obtain the regeneration of structure 1530. Structure 1530 is then subjected to several washes (by centrifugation) to remove the second bead 1516. The single-stranded complement 1512 bound to structure 1530 is a single-stranded barcode adapter.
[0248] As shown in Figure 15e, a fork-type adapter can be generated using single-stranded complement 1512. Next, to generate the fork-type adapter 1550, single-stranded complement 1512 is released from structure 1530, optionally using light, and then mixed with the universal complementary sequence 1526. To generate ligationable ends, a single-base T overhang is generated on the universal complementary sequence 1526 using a type II restriction enzyme. The T overhang fits with the A overhang on the polynucleotide target to be sequenced (not shown).
[0249] (Example 10) Production of fork-type adapter template barcode sequences and derived adapters by bead emulsion PCR As shown in Figure 16, structure 1600, comprising magnetic beads (1601) bound to a single-strand adapter-barcode sequence 1602, is produced according to the method described in Examples 8 and 9, or any other method described herein. Structure 1600 is then partitioned into a capsule (or another emulsion) 1620 by interfacial polymerization, for example, according to the method described herein. Capsule 1620 also contains reagents (e.g., polymerase, primers, dNTPs, buffer, salt) for amplification of the single-strand adapter-barcode sequence 1602 by asymmetric PCR. The reverse primer is present in excess of the forward primer, or vice versa, enabling asymmetric amplification. The single-strand adapter-barcode sequence 1602 is amplified 1603, and the reaction proceeds through the linear phase of amplification 1604 to produce a single-strand adapter product 1605 complementary to the single-strand barcode adapter-template 1602. At this point, capsule 1620 contains both a single-strand adapter 1605 in solution and a single-strand adapter-barcode sequence 1602 bound to a magnetic bead (1601). Capsule 1620 is then separated by magnetic separator 1606 from the part that does not contain the beads (and therefore the template 1602 and single-strand adapter 1605). Capsule 1620 can be ruptured to produce a fork-shaped adapter as described in Example 9.
[0250] (Example 11) Barcoding using bead emulsion PCR and fragmentation using fragmentase As shown in Figure 17, structure 1700, which includes a single-strand adapter-barcode sequence 1702 bonded to a magnetic bead (1701), is produced according to the method described in Examples 8 and 9, or any other method described herein. Interfacial polymerization is carried out on a droplet containing structure 1700, generating a capsule 1704 containing the single-strand adapter-barcode sequence 1702 bonded to the bead 1701 via a photounstable linker.
[0251] Two mixtures are prepared. Mixture Z1 contains a target polynucleotide (i.e., the polynucleotide to be fragmented and barcoded), a fragmentation enzyme (e.g., NEBNEXT DSDNA FRAGMENTASE), and a partially complementary universal sequence. The second mixture Z2 contains the capsule 1704 produced as described above, and magnesium chloride in a concentration sufficient to activate the fragmentation enzyme. Mixtures Z1, Z2, or both Z1 and Z2 also contain T4 polymerase, Taq polymerase, and a heat-stable ligase.
[0252] Mixtures Z1 and Z2 are mixed and allowed to form an intracapsule according to methods described elsewhere in this disclosure, such as flow focusing. Figure 17 illustrates an intracapsule produced according to the above method. The outer capsule 1703 contains the inner capsule 1704 and the medium 1705. The inner capsule 1704 is one member of a library of encapsulated bead-bound single-stranded barcode adapters. Thus, the inner capsule 1704 contains multiple copies of structure 1700, which can be used to generate a free single-stranded adapter-barcode sequence 1702, and the same barcode adapter can be bound to an intrapartition polynucleotide such as the outer capsule 1703.
[0253] Medium 1705 contains the contents of the above mixtures Z1 and Z2. More specifically, Medium 1705 contains a target polynucleotide 1706, a partially complementary universal sequence 1707, and an enzyme mix 1708 comprising a fragmentase, T4 polymerase, Taq polymerase, a heat-unstable ligase, magnesium chloride, and appropriate buffers.
[0254] During the formation of the capsule-in-capsule and exposure of the capsule-in-capsule to appropriate conditions, the enzymes process the target polynucleotide. More specifically, fragmentase fragments the target polynucleotide, and T4 polymerase blunts the ends of the fragmented target polynucleotide. The fragmentase and T4 polymerase are then thermally inactivated, and the inner capsule 1704 is ruptured using stimulation, releasing its contents into the outer capsule 1703. Taq polymerase adds 3'-A protrusions to the fragmented, blunt-terminated target polynucleotide. The single-stranded adapter-barcode sequence 1702 hybridizes with the partially complementary universal sequence 1707, is released from the bead by light, and forms a fork-shaped adapter having a 3'-T protrusion that fits with the 3'-A protrusion on the fragmented target polynucleotide. A thermally stable ligase ligates the fork-shaped adapter to the fragmented target polynucleotide, producing a barcoded target polynucleotide. Next, the outer capsule 1703 is ruptured, and the sample from all outer capsules is pooled and the target polynucleotide is sequenced. If necessary, additional preparation steps (e.g., bulk amplification, size selection, etc.) may be performed before sequencing.
[0255] In some cases, Z1 may contain multiple versions of the partially complementary universal sequence 1707. Furthermore, although this embodiment demonstrates barcoding of the target polynucleotide using a heat-stable ligase, this process can also be achieved using PCR.
[0256] (Example 12) Barcoding using bead emulsion PCR and fragmentation by sonication As shown in Figure 18, structure 1800, which includes a single-stranded adapter-barcode sequence 1802 bound to a magnetic bead (1801), is produced according to the method described in Examples 8 and 9, or any other method described herein. Interfacial polymerization is carried out on a droplet containing structure 1800 to produce capsule 1803 containing the single-stranded adapter-barcode sequence 1802 bound to the bead 1801 via a photounstable linker. The target polynucleotide (i.e., the polynucleotide to be fragmented) is fragmented into capsule 1804. Capsule 1804 containing the target polynucleotide is configured to withstand ultrasonic stress. Capsule 1804 containing the target polynucleotide is exposed to ultrasonic stress (e.g., COVARIS Focused-Ultrasonicator) to fragment the target polynucleotide and produce a fragmented target polynucleotide capsule.
[0257] A mixture Z1 is prepared containing capsule 1803, a fragmented target polynucleotide capsule 1804, a partially complementary universal sequence 1805, an enzyme mixture (T4 polymerase, Taq polymerase, and a heat-stable ligase) 1806, and a suitable buffer. The capsule-in-capsule is generated according to the method described elsewhere in this disclosure, such as flow focusing.
[0258] Figure 18 illustrates an inner capsule produced according to the method described above. The outer capsule 1807 comprises capsules 1803 and 1804 and a medium 1808. The inner capsules 1803 and 1804 each contain a capsule containing structure 1800 and a capsule containing fragmented target polynucleotide 1809, respectively. The inner capsule 1803 contains multiple copies of structure 1800, which can be used to generate a free single-stranded barcode adapter 1802, and the same barcode adapter can be bound to a polynucleotide in a partition, such as the fragmented polynucleotide 1809 contained in the inner capsule 1804.
[0259] Medium 1808 contains the contents of the above mixture Z1. More specifically, Medium 1808 contains a partially complementary universal sequence 1805, an enzyme mixture (T4 polymerase, Taq polymerase, and a heat-stable ligase) 1806, and a suitable buffer.
[0260] The inner capsule 1804, containing the fragmented target polynucleotide 1809, is exposed to a stimulus to rupture, releasing its contents into the contents of the outer capsule 1807. T4 polymerase blunts the ends of the fragmented target polynucleotide; Taq polymerase adds a 3'-A protrusion to the fragmented blunt-ended target polynucleotide. The T4 polymerase and Taq polymerase are then thermally inactivated, and a stimulus is applied to release the contents of the inner capsule 1803 into the outer capsule 1807. The single-stranded adapter-barcode sequence 1802 hybridizes with the partially complementary universal sequence 1805, and the adapter is released from the beads by light, forming a fork-shaped adapter having a 3'-T protrusion that fits with the 3'-A protrusion on the fragmented target polynucleotide. A thermally stable ligase ligates the fork-shaped adapter to the fragmented target polynucleotide, producing a barcoded target polynucleotide. Next, the outer capsule 1807 is ruptured, and the samples from all outer capsules are pooled and the target polynucleotide is sequenced.
[0261] In some cases, Z1 may contain multiple versions of the partially complementary universal sequence 1807. Furthermore, although this embodiment demonstrates barcoding of the target polynucleotide using a heat-stable ligase, this process can also be achieved using PCR.
[0262] (Example 13) Barcoding using multi-annealing and looping-based amplification (MALBAC) A primer containing Sequence ID No. 22 is prepared as shown in Figure 19a. This primer contains a barcode region (named "Barcode"), a primer sequencing region (named "PrimingSeq"), and an 8-nucleotide variable region (named "NNNNNNNN") which may contain any combination of A, T, C, or G. The primer shown in Figure 19 is mixed with the target polynucleotide (shown as a loop in Figure 19) together with a polymerase having chain displacement activity (e.g., Vent, exo+DeepVent, exo-DeepVent) in a partition (e.g., a capsule, an emulsion droplet, etc.). In some cases, a non-chain displacement polymerase (e.g., Taq, PfuUltra) is used. The partition is then subjected to MALBAC amplification. Suitable MALBAC cycling conditions are known and, for example, described in Zong et al., Science, 338(6114), pp. 1622–1626 (2012), the whole of which is incorporated herein by reference.
[0263] The looped MALBAC product is produced as shown in Figure 19b, as Sequence ID No. 23. The looped MALBAC product contains the original primer shown in Figure 19a, the target polynucleotide to be barcoded toward the loop, and a region complementary to and hybridized with the original primer sequence. The partitions are destroyed and their contents are recovered. In some cases, multiple partitions are generated. The partitions are destroyed collectively, and after recovering their contents, they are pooled.
[0264] Next, the generated MALBAC product shown in Figure 19b is treated with a restriction enzyme (e.g., BfuCl or similar) to generate a 4-base pair overhang on the MALBAC product (in this case, GATC shown in italics). This structure is represented by Sequence ID No. 24 and is shown in Figure 19c. The fork-shaped adapter shown as Sequence ID No. 25 in Figure 19d contains a complementary overhang (in this case, CTAG shown in bold) to the overhang generated on the MALBAC product. The fork-shaped adapter is mixed with the MALBAC product in Figure 19c, and the complementary regions hybridize. The fork-shaped adapter and the MALBAC product are ligated together using a thermostable ligase to form the desired structure, Figure 19e, as Sequence ID No. 26. Further regions (e.g., immobilization regions, further barcodes, etc.) can be added to the fork-shaped adapter using further amplification methods (e.g., PCR).
[0265] In some cases, other base pair protrusions (e.g., 1-base pair to 10-base pair protrusions) may be desirable. These protrusions can be generated using restriction enzymes and, if necessary, used as substitutes, such as those described herein. For example, a 2-base pair protrusion can be generated using Taq α It is generated on MALBAC products using I.
[0266] Alternatively, a primer can be designed as shown in Figure 19a, in which the RNA primer sequence is placed on the 5' side of the barcode region and a protrusion is generated using RNAase. As shown in Figure 19f, the MALBAC product 1900 contains an RNA primer sequence 1901 placed on the 5' side of the barcode region 1902. The MALBAC product 1900 also contains a sequencing primer region 1903, a target polynucleotide 1904, a complementary sequencing primer region 1905, a complementary barcode region 1906, and a region 1907 complementary to the RNA primer sequence 1901. The MALBAC product 1900 is treated with RNAse H 1908 to digest the RNA primer region sequence 1901 and obtain a 2-6 base pair protrusion 1909 on the MALBAC product 1900, thereby obtaining structure 1920. A universal complementary region 1910 containing a region complementary to the protrusion on structure 1910 is then added to structure 1910. Next, the universal complementary region 1910 hybridizes with structure 1920, and the universal complementary region 1910 is ligated to structure 1920 using a thermally stable ligase.
[0267] (Example 14) Barcoding using multi-annealing and looping-based amplification (MALBAC) As shown in Figure 20, template 2000, including the barcode region, is mixed with reagent 2001 required for PCR in capsule 2002, for example, by interfacial polymerization or any other method described herein. MALBAC primers are generated from template 2000 using PCR. Capsule 2000 is then encapsulated in an outer capsule 2003, which also contains a mixture 2004 containing the target polynucleotide 2005 to be barcoded and reagents 2006 required for MALBAC amplification (e.g., DeepVent polymerase, dNTP, buffer). Capsule 2002 is ruptured upon appropriate exposure of capsule 2002 to a stimulus designed to rupture it, and the contents of capsule 2002 mix with those of mixture 2004. MALBAC amplification of the target polynucleotide 2005 is initiated, producing a MALBAC product similar to that described as 1900 in Figure 19f.
[0268] Next, the outer capsule 2003 is ruptured using an appropriate stimulus and the contents are recovered. The MALBAC product is then treated with an appropriate restriction enzyme and coupled to the fork-shaped adapter in the substance described in Example 13. Further downstream preparation steps (e.g., bulk amplification, size selection, etc.) are then carried out as necessary.
[0269] (Example 15) Barcoding using multi-annealing and looping-based amplification (MALBAC) As shown in Figure 21a, MALBAC primer 2100 is prepared. MALBAC primer 2100 contains a sequence priming region 2101 and an 8-nucleotide variable region 2102. Primer 2100 is mixed with target polynucleotide 2103 together with a polymerase having chain displacement activity (e.g., Vent, exo+DeepVent, exo-DeepVent) in a partition (e.g., capsule, emulsion). In some cases, a non-chain displacement polymerase (e.g., Taq, PfuUltra) is used. The partition is then subjected to MALBAC amplification 2104.
[0270] A looped MALBAC product 2110 is produced, containing a sequencing priming region 2101, a target polynucleotide 2103, and a complementary sequencing priming region 2105. The MALBAC product 2110, shown in linear form 2120 in Figure 21b, is then contacted with another primer 2130 containing a sequencing primer region 2106, a barcode region 2107, and an immobilization region 2108. Primer 2130 is produced using asymmetric digital PCR. Using a single cycle of PCR, the primer is used to produce a double-stranded product 2140 containing primer 2130 and therefore the barcode region 2107.
[0271] Next, the double-stranded product 2140 is denatured and then brought into contact with another primer 2150 shown in Figure 21c. Primer 2150 includes a barcode region 2109, a sequencing primer region 2111, and an immobilization region 2112. In the presence of primers 2113 and 2114, further rounds of PCR can add the barcode region 2109 to the end of the target polynucleotide bound to the barcode region 2107. Further downstream preparation steps (e.g., bulk amplification, size selection) are then performed as needed.
[0272] (Example 16) Barcoding using multi-annealing and looping-based amplification (MALBAC) As shown in Figure 22, the primer template 2200 containing the barcode region is mixed with reagent 2201 required for PCR in capsule 2202, for example, by interfacial polymerization or any other method described herein. Then, primers are generated from the template 2200 using PCR. Next, capsule 2202 is encapsulated in an outer capsule 2203 which also contains a mixture 2204 containing the target polynucleotide 2205 to be barcoded, reagent 2206 required for MALBAC amplification (e.g., DeepVent polymerase, dNTP, buffer), and a MALBAC primer 2207 without a barcode (similar to the MALBAC primer 2100 described in Example 15). MALBAC amplification of the target polynucleotide 2205 is initiated, producing a MALBAC product similar to that described as 2110 in Figure 21a. Next, capsule 2202 is destroyed upon proper exposure to a stimulus designed to rupture capsule 2202, and the contents of capsule 2202 are mixed with those of mixture 2204. A single cycle of PCR is initiated using primers generated from template 2200 to produce a barcoded product similar to that described in Example 15.
[0273] Next, the outer capsule 2203 is ruptured using an appropriate stimulus, and the contents are recovered. Then, further downstream preparation steps (e.g., bulk amplification, size selection, addition of further barcodes, etc.) are carried out as needed.
[0274] (Example 17) Barcoding using transposases and tagmentation As shown in Figure 23, the single-stranded adapter-barcode polynucleotide sequence 2300 is synthesized, segmented, amplified, and sorted as described in Example 1 or by any other method described herein. Interfacial polymerization is carried out on droplets containing the single-stranded adapter-barcode polynucleotide sequence to produce capsules 2301.
[0275] Two mixtures are prepared. Mixture Z1 comprises the target polynucleotide 2302 (i.e., the polynucleotide to be fragmented and barcoded), the transposome 2303, and the partially complementary universal sequence 2304. The second mixture Z2 comprises the capsule 2301 produced as described above and the reagent 2305 required for PCR as described elsewhere herein.
[0276] Mixtures Z1 and Z2 are mixed and allowed to form an intracapsule according to a method described elsewhere in this disclosure, such as flow focusing. Figure 23 illustrates an intracapsule produced according to the above method. The outer capsule 2306 contains capsule 2301 and medium 2307. Capsule 2301 is one member of a library of encapsulated single-stranded adapter-barcode polynucleotides. Thus, capsule 2301 contains multiple copies of the single-stranded adapter-barcode polynucleotide sequence 2300, which can be used to encapsulate the same barcode onto a polynucleotide in a partition such as the outer capsule 2306.
[0277] Medium 2307 contains the contents of the above mixtures Z1 and Z2. More specifically, Medium 2307 contains reagents 2305 necessary for PCR, such as the target polynucleotide 2302, the partially complementary universal sequence 2304, and the hot-start Taq.
[0278] During the formation of the capsule-in-a-capsule and exposure of the capsule-in-a-capsule to appropriate conditions, the transposome processes the target polynucleotide. More specifically, the transposase fragments the target polynucleotide via tagmentation and tags it with a common priming sequence. The tagged target polynucleotide is then heated to fill any gaps in the target polynucleotide generated by the transposase. The transposase is then thermally inactivated, and the inner capsule 2301 is ruptured using a stimulus, releasing its contents into the outer capsule 2306. The outer capsule 2306 is heated to 95°C to activate the hot-start Taq. The reaction is carried out using limited-cycle PCR to add the single-stranded adapter-barcode polynucleotide sequence 2300 to the target polynucleotide 2302. The outer capsule 2306 is then ruptured, and the target polynucleotide is sequenced.
[0279] While specific embodiments have been illustrated and described, various modifications can be made thereto, and it should be understood from the above that this is intended. Furthermore, the present invention is not intended to be limited by the specific examples provided in the specification. Although the present invention has been described with reference to the specification above, the description and illustration of preferred embodiments in this specification is not meant to be constrained. Moreover, it should be understood that all aspects of the present invention are not limited to the specific expressions, configurations or relative proportions described herein, which depend on various conditions and variables. Various modifications in the forms and details of embodiments of the present invention will be apparent to those skilled in the art. Therefore, the present invention is intended to encompass any such modifications, changes and equivalents. The following claims define the scope of the present invention, and the methods and structures within these claims and their equivalents are intended to be encompassed thereby. [Explanation of Symbols]
[0280] 101 First immobilization region 102 Second immobilization region 103 First sequencing primer region 104 Second sequencing primer region 105 Target polynucleotides 106 Fork-type adapter structure 201 First immobilization region 202 Second immobilization region 203 First sequencing primer region 204 Second sequencing primer region 205 Barcode (BC1) 206 Barcode (BC2) 207 Barcode (BC3) 401 Single-stranded adapter-barcode polynucleotide sequence 402 First immobilization region 403 Barcode area 404 First sequencing primer region 405 Double-stranded products 406 Single-chain products 407 Partially Complementary Universal Sequences 408 Second immobilization region 409 Second sequencing primer region 410 Exponential period of amplification 411 Linear period of amplification 412 Fluorescence-Assisted Cell Sorting System (FACS) 413 Partially annealed fork-shaped structure 501 Outer capsule 502 Inner capsule 503 Multiple Copies Single-Stranded Adapter - Barcode Polynucleotide 504 Medium 505 Target Polynucleotides 506 Partially Complementary Universal Sequences 507 Enzyme Mix 601 Outer capsule 602 Inner capsule 603 Single-stranded adapter - barcode polynucleotide 604 Medium 605 Inner capsule 606 Fragmented Target Polynucleotides 607 Partially Complementary Universal Sequences 608 Enzyme mixture 701 Hairpin Adapter 702 Double-stranded region 703 3'-T protrusion for AT ligation 704 Regions that can be cleaved by restriction enzymes 1101 First single-stranded adapter-barcode polynucleotide sequence 1102 First immobilization region 1103 First barcode area 1104 First sequencing primer region 1105 Double-stranded products 1106 Single-chain products 1106 Multiple Copy Single-Stranded Adapter - Barcode Polynucleotide 1110 Exponential period of amplification 1111 Linear period of amplification 1112 Fluorescence-Assisted Cell Sorting System (FACS) 1120 capsules 1131 Second single-stranded adapter-barcode polynucleotide sequence 1132 Second immobilization region 1133 Second barcode area 1134 Second sequencing primer region 1135 Double-stranded products 1136 Single-chain products 1136 Multiple Copy Single-Stranded Adapter - Barcode Polynucleotide 1150 capsules 1160 Outer capsule 1170 Target Polynucleotides 1180 capsules 1180 Enzyme Mix 1190 Medium 1201 The first liquid 1202 Second liquid 1203 T connection 1204 The capsule outer shell begins to form. 1301 The first liquid 1302 Second liquid 1303 T connection 1304 The second capsule shell begins to form. 1401 Single-strand adapter - barcode array 1402 First immobilization region 1403 Barcode area 1404 First sequencing primer region 1405 The first bead 1406 RNA primer 1407 Amplify polynucleotides 1408 Double-stranded product 1409 Cleaning 1410 Degeneration 1411 Cleaning 1412 Single-stranded complement 1413 Complementary immobilization region 1414 Complementary barcode area 1415 Complementary sequencing primer region 1416 The second bead 1417 DNA polynucleotides 1418 Centrifugal Separation 1419 NaOH 1420 Structure 1424 Release 1425 Mixed 1426 Universal Complementary Sequence 1430 Structure 1440 Structure 1450 Fork-type adapter 1501 Single-strand adapter - barcode array 1502 First immobilization region 1503 Barcode area 1504 First sequencing primer region 1505 First bead 1506 RNA primer 1507 Amplifying polynucleotides 1508 Double-stranded product 1509 Cleaning 1510 Sodium hydroxide (NaOH) 1511 Cleaning 1512 Single-stranded complement 1513 Complementary immobilization region 1514 Complementary barcode area 1515 Complementary sequencing primer region 1516 The second bead 1517 Complementary DNA polynucleotides 1518 Centrifugal Separation 1519 NaOH 1520 Structure 1524 Release 1525 Mixed 1526 Universal Complementary Sequence 1530 Structure 1540 Structure 1550 Fork-type adapter 1600 Structure 1601 Magnetic Beads 1602 Single-strand adapter - barcode array 1603 Amplification 1604 Linear period of amplification 1605 Single-strand adapter, single-strand adapter product 1606 Magnetic separation 1620 capsules 1700 Structure 1701 Magnetic Beads 1702 Single-strand adapter - barcode array 1703 Outer capsule 1704 Inner capsule 1705 Medium 1706 Target polynucleotides 1707 Partially complementary universal sequence 1708 Enzyme Mix 1800 Structure 1801 Magnetic Beads 1802 Single-strand adapter - barcode array 1803 Inner capsule 1804 Inner capsule, targeted polynucleotide capsule 1805 Partially Complementary Universal Sequence 1806 Enzyme mixture 1807 Outer capsule 1808 Medium 1809 Fragmented target polynucleotides 1900 MALBAC products 1901 RNA primer sequence 1902 Barcode area 1903 Sequencing primer region 1904 Target polynucleotide 1905 Complementary sequencing primer region 1906 Complementary barcode area 1907 RNA primer sequence and complementary region 1908 RNAse H 1909 Base pair protrusion 1910 Structure 1910 Universal Complementary Area 1920 Structure 2000 mold 2001 Pharmaceuticals 2002 Capsules 2003 Outer capsule 2004 mixture 2005 Targeted Polynucleotides 2006 Reagents 2100 MALBAC Primer 2101 Sequence priming region 2102 8-nucleotide variable region 2103 Target polynucleotides 2104 MALBAC amplification 2105 Complementary sequence priming region 2106 Sequencing primer region 2107 Barcode area 2108 Immobilization area 2109 Barcode area 2110 MALBAC products 2111 Sequencing primer region 2112 Immobilization area 2113 Primer 2114 Primer 2120 Linear form 2130 Primer 2140 Double-stranded products 2150 Primer 2200 mold 2201 Reagent 2202 capsules 2203 Outer capsule 2204 Mixture 2205 Target polynucleotide 2206 Reagents 2207 MALBAC Primer 2300 Single-stranded adapter - barcode polynucleotide sequence 2301 capsules, inner capsule 2302 Target polynucleotides 2303 Transpososome 2304 Partially Complementary Universal Sequence 2305 Reagent 2306 Outer capsule 2307 Medium
Claims
[Claim 1] A barcode library as described herein.