Methods of barcoded nucleic acids for detection and sequencing
By using barcode templates and sequencing techniques at the single-cell level, the problem of difficulty in single-cell genotyping in the prior art is solved, and high-sensitivity cell classification and somatic mutation detection are achieved.
Patent Information
- Application Number
- CN202380076657.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-08-29
- Filing Date
- 2023-08-29
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to genotypify at the single-cell level, especially when distinguishing tumor cells from normal cells, and most sequencing methods do not provide complete haplotype phased information.
By isolating multiple cells or nuclei into the compartment, amplifying and attaching the barcode template to the cell content fragments, sequencing is performed to classify fragments belonging to the same cell unit.
A method of characterizing biological samples at the single-cell level is realized, which can effectively distinguish the genomic information of different cells, eliminate normal cell background noise, and improve the sensitivity of somatic mutation detection.
Smart Images

Figure CN120153092A_ABST
Abstract
Description
Cross - Reference to Related Applications
[0001] This application claims the benefit and priority of U.S. Provisional Application No. 63 / 373,778, filed Aug. 29, 2022, the entire disclosure of which is incorporated herein by reference. All publications, patents, and other documents mentioned herein are incorporated by reference in their entireties. Technical Field
[0002] The present disclosure pertains to the technical field of genomics. More specifically, the present disclosure generally relates to nucleic acid sequencing. More specifically, the present disclosure relates to methods for improving nucleic acid detection and sequencing for single - cell analysis, haplotype phasing, de novo assembly, and variant detection. Background Art
[0003] Nucleic acid sequencing can provide information for a variety of biomedical applications, including diagnosis, prognosis, pharmacogenomics, and forensic biology. Sequencing may involve basic low - throughput methods, including Maxam - Gilbert sequencing (chemically modified nucleotides) and Sanger sequencing (chain termination), or high - throughput next - generation methods, including massively parallel pyrosequencing, sequencing - by - synthesis, sequencing - by - ligation, semiconductor sequencing, and others. For most sequencing methods, the sample (such as a nucleic acid target) needs to be processed before being introduced into the sequencing instrument. For example, the sample can be fragmented, amplified, or attached to an identifier. Unique identifiers are typically used to identify the origin of a particular sample. Most sequencing methods produce relatively short sequencing reads, ranging from dozens to hundreds of bases in length, and due to the sequencing read length limitation, cannot generate complete haplotype phase information. Most biological samples contain many cells. And most assays measure responses to bulk cells rather than at the single - cell level. There is a need in the art for methods of genotyping cells at the single - cell level, for example, to separate tumor cells from wild - type or normal cells in a sample. Such methods are provided by the methods and features described herein. Summary of the Invention
[0004] The present disclosure provides methods for improved nucleic acid detection and sequencing. In particular, the present disclosure provides improved methods for single - cell nucleic acid sequencing and detection.
[0005] On the one hand, the present disclosure provides a single-cell sequencing method for characterizing a biological sample at the single-cell level. The method includes isolating a plurality of cells or a plurality of cell nuclei into compartments, wherein each cell or cell nucleus is isolated into a separate compartment having a plurality of barcode templates, wherein each barcode template contains a barcode sequence, and wherein at least some of the compartments contain more than one group of barcode templates, each group of barcode templates having a unique barcode sequence different from other groups of barcode templates. The method includes amplifying at least one type of cell content in each of the individual cells or cell nuclei into a plurality of copies, and fragmenting the cell content in each of the compartments into a plurality of fragments. The method includes attaching the barcode templates to the individual fragments. The method includes collecting the fragments to which the barcode templates are attached. The method includes sequencing the fragments to which the barcodes are attached and classifying the fragments having the same barcode sequence as belonging to the same cell unit.
[0006] On the one hand, the present disclosure provides a single-cell sequencing method for characterizing a biological sample at the single-cell level. The method includes isolating a plurality of cells or a plurality of cell nuclei and a plurality of barcode templates into compartments, wherein each cell or cell nucleus is isolated into a separate compartment having at least one barcode template containing a barcode sequence, and wherein at least some of the compartments contain at least two different barcode templates, each different barcode template having a different barcode sequence. The method includes amplifying at least one type of cell content in each of the individual cells or cell nuclei into a plurality of copies, fragmenting the cell content in each of the compartments into fragments, and amplifying the at least one barcode template in each of the compartments. The method includes attaching the barcode templates to the individual fragments. The method includes collecting the fragments to which the barcode templates are attached. The method includes sequencing the fragments to which the barcodes are attached and classifying the fragments having the same barcode sequence as belonging to the same cell unit.
[0007] On the one hand, the present disclosure provides a method for single-cell transcriptome sequencing. The method includes generating cDNA from cellular RNA or nuclear RNA of a single cell or cell nucleus among a plurality of cells or cell nuclei. The method includes randomly tagging the generated cDNA along the full length of the cDNA in each of the individual cells or cell nuclei using a plurality of transposomes to form a plurality of tagged cDNA fragments, wherein each transposome contains at least one transposon and a transposase. The method includes isolating the plurality of cells or cell nuclei into compartments, wherein each cell or cell nucleus is isolated into a separate compartment having a plurality of barcode templates, wherein each barcode template contains a barcode sequence. The method includes attaching the barcode templates to each of the tagged cDNA fragments in the compartments. The method includes collecting the tagged cDNA fragments to which the barcodes are attached. The method includes sequencing the barcodes and the tagged cDNA fragments to which the barcodes are attached to characterize the transcriptome profile of each of the individual cells or cell nuclei on a single-cell basis.
[0008] On the one hand, the present disclosure provides a method for single-cell transcriptome sequencing. The method includes generating cDNA from cellular RNA or nuclear RNA of a single cell or nucleus in a plurality of cells or nuclei. The method includes randomly tagging the generated cDNA along the full length of the cDNA in each cell or nucleus to form a plurality of tagged cDNA fragments using a plurality of transposomes, wherein each transposome includes at least one transposon and a transposase. The method includes isolating the cells or nuclei from a plurality of barcode templates, wherein each cell or nucleus is isolated into a separate compartment having at least one barcode template. The method includes attaching the barcode template to each of the tagged cDNA fragments. The method includes collecting the barcode-attached cDNA fragments. The method includes sequencing the barcode and the barcode-attached cDNA fragments to characterize the transcriptome profile of each cell on a single-cell basis.
[0009] In any of the above aspects or embodiments thereof, each barcode template is a nucleotide sequence capable of functioning as a unique identifier.
[0010] In any of the above aspects or embodiments thereof, each barcode template is freely present in solution. In any of the above aspects or embodiments thereof, each barcode template is immobilized on a carrier. In any of the above aspects or embodiments thereof, the carrier is a solid bead or particle, a soluble bead or particle, or a combination thereof.
[0011] In any of the above aspects or embodiments thereof, the type of cellular content is RNA, DNA, RNA / DNA hybrid, protein, metabolite, ligand, compound, drug, macromolecule, or a combination thereof. In any of the above aspects or embodiments thereof, the type of cellular content is RNA, DNA, RNA / DNA hybrid, or a combination thereof.
[0012] In any of the above aspects or embodiments thereof, the fragment is directly attached to the barcode template. In any of the above aspects or embodiments thereof, the fragment is indirectly attached to the barcode template. In any of the above aspects or embodiments thereof, the fragment is attached to an adaptor oligomer or linker, wherein the adaptor oligomer or linker is attached to the barcode template.
[0013] In any of the above aspects or embodiments thereof, the cellular content is endogenous. In any of the above aspects or embodiments thereof, the cellular content is exogenous.
[0014] In any of the above aspects or their embodiments, the compartment includes cells or cell nuclei that are not further compartmentalized; tubes or microtubes; pores or micropores; plates; wells in a multi-well plate; slides; spots on a slide; droplets; conduits; channels; bottles; chambers; or flow cells.
[0015] In any of the above aspects or their embodiments, the steps of amplifying the cell contents and / or the barcode templates and the step of attaching the barcode templates to the fragments are carried out substantially simultaneously.
[0016] In any of the above aspects or their embodiments, the method further includes: identifying barcode sequences attached to cell contents derived from the same cell or cell nucleus, and pooling cell units corresponding to barcode sequences identified as being attached to cell contents derived from the same cell or cell nucleus.
[0017] In any of the above aspects or their embodiments, the cells are eukaryotic cells, prokaryotic cells, or a combination thereof.
[0018] In any of the above aspects or their embodiments, the plurality of barcode templates in each compartment includes at least two barcode template groups, where each barcode template group has a different barcode sequence.
[0019] In any of the above aspects or their embodiments, the attachment results in at least two cDNA fragment groups, each cDNA fragment group being attached to a different barcode template group.
[0020] In any of the above aspects or their embodiments, the at least one barcode template is at least two different barcode templates, each barcode template having a different barcode sequence.
[0021] In any of the above aspects or their embodiments, the generated cDNA is first-strand cDNA and forms a DNA / RNA hybrid with cellular RNA or nuclear RNA.
[0022] In any of the above aspects or their embodiments, the generated cDNA is first-strand and second-strand cDNA and forms double-stranded DNA.
[0023] In any of the above aspects or their embodiments, the generated cDNA contains transcripts that contain the 3' and 5' ends of cellular RNA or nuclear RNA.
[0024] In any of the above aspects or their embodiments, the transcriptome profile includes the 3' and 5' ends of cellular RNA or nuclear RNA.
[0025] In any of the above aspects or their embodiments, the sequence of the cDNA fragment to which the barcode template is attached is converted into a full-length RNA sequence.
[0026] In any of the above aspects or their embodiments, attaching the barcode template to the tagged cDNA fragment includes amplifying the barcode template and / or amplifying the tagged cDNA fragment.
[0027] In any of the above aspects or their embodiments, the amplification of the barcode template and the amplification of the tagged cDNA fragment are performed separately.
[0028] In any of the above aspects or their embodiments, the amplification of the barcode template and the amplification of the tagged cDNA fragment are performed simultaneously.
[0029] In any of the above aspects or their embodiments, at least one barcode template in each compartment is a single barcode template.
[0030] In any of the above aspects or their embodiments, multiple barcode templates in each compartment are multiple copies of the same barcode template.
[0031] In any of the above aspects or their embodiments, the cell or nucleus, or the multiple cells or nuclei, are obtained from a biological sample or a cell culture. Further aspects
[0032] In one aspect, an amplifiable single-cell sequencing method for characterizing a biological sample at the single-cell level is described and provided. The method includes: providing a plurality of cells or cell nuclei from the sample, providing a plurality of barcode templates, isolating the cells or cell nuclei and more than one different barcode template in a compartment; amplifying each barcode template into a plurality of copies, and amplifying one or more types of cell contents into a plurality of copies, wherein the cell contents include naturally occurring nucleic acid sequences or nucleic acid sequences artificially attached to nucleic acid sequences, in the isolated compartment; coupling the amplified barcode templates with the amplified cell contents in the compartment; performing sequencing to determine the barcode sequences in the barcode templates and their associated cell content sequences; classifying the cell contents having the same barcode sequence as a cell unit or a part of a cell unit. In one embodiment, the amplification step and the coupling step can be performed sequentially or simultaneously. These methods turn one cell content into more than one cell unit. In an embodiment, the cell contents include DNA, RNA, proteins, lipids, or organelles inside the cell, or organelles in the cell nucleus, or organelles related to the outside of the cell or a combination thereof. In an embodiment, the cells are eukaryotic cells and / or prokaryotic cells. In an embodiment, the compartment is a well, a micro-well, a droplet, a micro-droplet, a pore, and other materials capable of physically isolating cell contents into different reaction units or spaces.
[0033] In one aspect, a method for sequencing the full-length transcriptome of single cells is provided, wherein the method includes: providing a plurality of cells from a biological sample; contacting the cells with a reverse transcriptase and a primer (such as an oligo-dT primer) to generate first-strand cDNA in situ; providing a plurality of transposomes, each transposome comprising at least one transposon and a transposase; randomly tagging RNA / cDNA hybrid transcripts in situ across the entire transcript; providing a plurality of barcode templates and amplification reagents; compartmentalizing the cells, barcode templates, and amplification reagents to produce two or more compartments, wherein each compartment contains a cell, one or more barcode templates having different barcode sequences, and amplification reagents; amplifying the barcode templates and the tagged RNA / cDNA fragments, attaching the barcode sequences to the cDNA fragments or fragments generated from the cDNA such that a plurality of barcode-attached fragments share the same one or more barcode sequences present in the compartment; collecting the barcode-attached fragments; sequencing the barcodes and the barcoded and tagged nucleic acids to characterize the full-length transcriptome profile on a single-cell basis.
[0034] On the other hand, a method for tracking the origin of a target by barcode tagging is provided. The method includes encapsulating at least one unique barcode template and at least one target in a compartment; amplifying one or more barcode templates and modifying the target, wherein the modified target is capable of linking to the barcode in the compartment; linking the barcode sequence to the modified target such that multiple modified targets share the same one or more barcode sequences present in the compartment; removing the compartment and collecting the modified targets with barcode tags for downstream applications. In one embodiment, the target is selected from nucleic acids, proteins including antibodies, ligands, compounds, cell nuclei, cells, and combinations thereof. In one embodiment, the cell can be a prokaryotic cell or a eukaryotic cell. In one embodiment, the modification of the target is selected from strand transfer reactions, tagging reactions, reverse transcription, amplification, primer extension, restriction digestion, hybridization, ligation, fragmentation, and combinations thereof. In some embodiments, the target is processed and / or modified prior to encapsulation. The processing methods are selected from the group consisting of denaturation, permeabilization, fixation, labeling, conjugation, in situ reactions, and combinations thereof. In some embodiments, the compartment origin of different barcode sequences present in the same compartment can be identified based on the compartment contents they share.
[0035] In some embodiments, the barcode template contains a central barcode sequence flanked by at least two handle sequences that can serve as primer sites, hybridization sites, or binding sites.
[0036] On the other hand, a method for tracking the origin of nucleic acid fragments by barcode tagging is provided. The method includes providing a plurality of nucleic acid targets and a plurality of transposomes, each transposome comprising at least one transposon and a transposase; incubating the nucleic acid targets and the transposomes together to form a strand transfer complex (STC) on the nucleic acid targets; providing a plurality of unique barcode templates; compartmentalizing the nucleic acid targets with the STC and the barcode templates to produce two or more compartments containing one or more nucleic acid targets and one or more than one barcode template with different barcode sequences; amplifying the barcode templates in the compartments, fragmenting the nucleic acid targets by disrupting the STC to form tagged nucleic acid fragments, and attaching the barcode sequences to the tagged nucleic acid fragments such that multiple fragments share the same one or more barcode sequences present in the compartment; removing the compartments and collecting the nucleic acid fragments with barcode tags.
[0037] On the other hand, a method for tracking the origin of nucleic acid fragments by barcoding plus tagging is provided. The method includes providing a plurality of nucleic acid targets and a plurality of transposomes, each transposome comprising at least one transposon and a transposase; incubating the nucleic acid targets and the transposomes together to form a strand transfer complex (STC) on the nucleic acid targets; providing a plurality of unique barcode templates; compartmentalizing the nucleic acid targets with the STC and the barcode templates to generate two or more compartments, which contain one or more nucleic acid targets and one or more barcode templates having different barcode sequences; attaching the barcode sequences to the nucleic acid targets in the compartments by: i) fragmenting the nucleic acid targets by disrupting the STC to form tagged nucleic acid fragments; ii) amplifying the tagged nucleic acid fragments with non-target-specific primers (i.e., only transposon-specific) and amplifying one or more barcode templates; iii) ligating the barcode templates to the tagged nucleic acid fragments, wherein the plurality of fragments share the same one or more barcode sequences present in the compartment; removing the compartments and collecting the nucleic acid fragments with barcode tags for downstream applications. For example, downstream applications include generating haplotype phasing sequencing information.
[0038] On the other hand, a method for tracking the origin of targeted nucleic acid fragments by barcoding plus tagging is provided. The method includes providing a plurality of nucleic acid targets, a plurality of target-specific primers, and a plurality of transposomes, each transposome comprising at least one transposon and a transposase; incubating the nucleic acid targets and the transposomes together to form a strand transfer complex (STC) on the nucleic acid targets; providing a plurality of unique barcode templates; compartmentalizing the nucleic acid targets with the STC and the barcode templates to generate two or more compartments, which contain one or more nucleic acid targets and one or more barcode templates having different barcode sequences; attaching the barcode sequences to the nucleic acid targets in the compartments by: i) fragmenting the nucleic acid targets by disrupting the STC to form tagged nucleic acid fragments; ii) amplifying the tagged nucleic acid fragments with transposon-specific primers and target-specific primers and amplifying one or more barcode templates; iii) ligating the barcode templates to the tagged nucleic acid fragments, wherein the plurality of fragments share the same one or more barcode sequences present in the compartment; removing the compartments; and collecting the nucleic acid fragments with barcode tags. In some embodiments, the nucleic acid targets are within a cell or nucleus, wherein the cell or nucleus is permeabilized or fixed and then incubated with the plurality of transposomes before compartmentalization with the target-specific primers and barcode templates.
[0039] On the other hand, a method for tracking the source of targeted nucleic acid fragments by barcode plus tag is provided. The method includes providing a plurality of nucleic acid fragments, a plurality of unique barcode templates, and a plurality of target-specific primers, wherein at least some of the target-specific primers are capable of attaching directly or indirectly to the barcode template; compartmentalizing the nucleic acid fragments, the target-specific primers, and the barcode template to produce two or more compartments, which contain one or more nucleic acid fragments, target-specific primers, and one or more barcode templates having different barcode sequences; attaching a barcode sequence to the nucleic acid fragments in the compartment by: i) amplifying the target from the nucleic acid fragments using the target-specific primers and amplifying one or more barcode templates; ii) ligating the barcode template to the amplified nucleic acid target in the compartment, wherein the plurality of amplified nucleic acid targets share the same one or more barcode sequences present in the compartment; iii) removing the compartment; and iv) collecting the barcoded nucleic acid targets for further analysis including sequencing.
[0040] On the one hand, a method for single-cell ATAC-seq is provided. The method includes providing a plurality of cells or cell nuclei and a plurality of transposomes, wherein each transposome contains at least one transposon and a transposase; incubating the plurality of cells or cell nuclei together with the plurality of transposomes to form strand transfer complexes (STCs) on the accessible chromatin in the cells or cell nuclei; providing a plurality of unique barcode templates; compartmentalizing the processed cells or cell nuclei and the barcode templates to produce two or more compartments, which include the cells or cell nuclei and one or more barcode templates having different barcode sequences; amplifying the barcode templates in the compartments, breaking the cell membrane and / or nuclear membrane, fragmenting the accessible chromatin by disrupting the STCs to form tagged nucleic acid fragments, and attaching the barcode sequence to the tagged nucleic acid fragments such that the plurality of fragments share the same one or more barcode sequences present in the compartment; removing the compartments and collecting the barcoded nucleic acid fragments; sequencing the barcodes and the barcoded nucleic acid to characterize the accessible chromatin regions on a single-cell basis.
[0041] In one aspect, a method for single-cell ATAC-seq is provided, the method comprising providing a plurality of cells or cell nuclei and a plurality of transposomes, wherein each transposome comprises at least one transposon and a transposase; incubating the plurality of cells or cell nuclei and the plurality of transposomes together to form a strand transfer complex (STC) on the accessible chromatin in the cell nuclei; providing a plurality of unique barcode templates; compartmentalizing the processed cells or cell nuclei and the barcode templates to produce two or more compartments, the compartments comprising cells or cell nuclei and one or more than one barcode templates having different barcode sequences; attaching the barcode sequences to the accessible chromatin fragments in the compartments by: i) breaking the cell membrane and / or nuclear membrane and fragmenting the accessible chromatin by disrupting the STC to form tagged nucleic acid fragments; ii) amplifying the tagged nucleic acid fragments and amplifying the barcode templates; iii) ligating the barcode templates to the tagged nucleic acid fragments, wherein a plurality of fragments share the same one or more barcode sequences present in the compartment; removing the compartments and collecting the barcoded nucleic acid fragments; sequencing the barcodes and the barcoded nucleic acids to characterize the accessible chromatin regions on a single-cell basis.
[0042] In one aspect, a method for barcoding single-cell whole genomes is provided. The method comprises providing a plurality of cells or cell nuclei and fixing the cells or cell nuclei to dissociate DNA from the proteins within the cells or cell nuclei; providing a plurality of transposomes, wherein each transposome comprises at least one transposon and a transposase; incubating the fixed cells or cell nuclei and the transposomes to form a strand transfer complex (STC) on the DNA within the fixed cells or cell nuclei; providing a plurality of unique barcode templates; compartmentalizing the processed cell nuclei and the barcode templates to produce two or more compartments, the compartments comprising cells or cell nuclei and one or more than one barcode templates having different barcode sequences; amplifying the barcode templates in the compartments, breaking the cell membrane and / or nuclear membrane, fragmenting the DNA by disrupting the STC to form tagged nucleic acid fragments; attaching the barcode sequences to the tagged nucleic acid fragments such that a plurality of fragments present in the compartment share one or more identical barcode sequences; removing the compartments; and collecting the barcoded nucleic acid fragments. In some embodiments, the strand transfer reaction occurs after compartmentalization of the cells or cell nuclei with one or more barcode templates. In embodiments, the cells are prokaryotic cells or eukaryotic cells.
[0043] In one aspect, a method for barcoding the whole genome of single cells is provided, wherein the method includes providing a plurality of cells or cell nuclei and fixing the cells or cell nuclei to dissociate DNA from the proteins within the cells or cell nuclei; providing a plurality of transposomes, wherein each transposome comprises at least one transposon and a transposase; incubating the fixed cells or cell nuclei and the transposomes to form strand transfer complexes (STCs) on the DNA within the fixed cells or cell nuclei; providing a plurality of unique barcode templates; compartmentalizing the treated cell nuclei and the barcode templates to produce two or more compartments, which include the cells or cell nuclei and one or more barcode templates having different barcode sequences; attaching barcode sequences to the genomic DNA in the cells or cell nuclei in the compartments by: i) disrupting the nuclear membrane and fragmenting the genomic DNA by disrupting the STCs to form tagged nucleic acid fragments; ii) amplifying the tagged nucleic acid fragments and amplifying the barcode templates; iii) ligating the barcode templates to the tagged nucleic acid fragments, wherein multiple fragments share the same one or more barcode sequences present in the compartment; removing the compartments and collecting the nucleic acid fragments with barcode tags. In some embodiments, the strand transfer reaction occurs after compartmentalization of the cells or cell nuclei with one or more barcode templates. In embodiments, the cells are prokaryotic cells or eukaryotic cells.
[0044] In one aspect, a method for single cell targeted sequencing is provided, the method including providing a plurality of cells and / or cell nuclei, providing a plurality of unique barcode templates, and providing a plurality of target-specific primers, wherein at least some of the target-specific primers are also capable of directly or indirectly attaching to the barcode templates; compartmentalizing the cells and / or cell nuclei, the barcode templates, and the target-specific primers to produce two or more compartments, the compartments including the cells and / or cell nuclei, one or more barcode templates having different barcode sequences, and the target-specific primers; amplifying the barcode templates in the compartments, ligating barcode sequences to the target-specific primers, disrupting the cell membrane / nuclear membrane, priming target genomic regions with the target-specific primers to generate barcode-attached target fragments such that multiple barcode-attached target fragments share the same one or more barcode sequences present in the compartment; removing the compartments and collecting the barcode-attached target fragments; and sequencing the barcodes and the barcoded nucleic acids to characterize the targeted regions on a per cell basis. In some embodiments, DNA or RNA is the target, or both DNA and RNA are targets. When RNA is the target, in addition to DNA polymerase, a reverse transcriptase is included.
[0045] In one aspect, methods for single cell targeted sequencing are provided, wherein the methods comprise: providing a plurality of cells and / or cell nuclei; providing a plurality of unique barcode templates; and providing a plurality of target-specific primers, wherein the target-specific primers are capable of directly or indirectly attaching to the barcode templates; compartmentalizing the cells and / or cell nuclei, wherein the barcode templates and target-specific primers create two or more compartments, the compartments containing cells and / or cell nuclei, one or more barcode templates with different barcode sequences, and target-specific primers; attaching barcode sequences to targeted nucleic acid fragments in the compartments by: i) disrupting the cell membrane and / or nuclear membrane to release nucleic acids; ii) amplifying the nucleic acid targets and amplifying the barcode templates; iii) ligating the barcode templates to the amplified nucleic acid targets, wherein a plurality of nucleic acid targets share the same one or more barcode sequences present in the compartment; removing the compartments and collecting the barcode-attached target fragments; and sequencing the barcodes and barcoded nucleic acids to characterize the targeted regions on a per cell basis. DNA or RNA is the target, or both DNA and RNA are targets. When RNA is the target, in addition to DNA polymerase, a reverse transcriptase is included.
[0046] In another aspect, methods for single cell RNA sequencing are provided, wherein the methods comprise providing a plurality of cells or cell nuclei, providing a plurality of unique barcode templates, providing a reverse transcriptase, and providing a plurality of primers that are capable of priming cDNA synthesis, or priming barcode template amplification, or priming cDNA, or a combination thereof; compartmentalizing the cells, barcode templates, reverse transcriptase, and primers to create two or more compartments, the compartments containing cells, one or more than one barcode templates with different barcode sequences, reverse transcriptase, and primers; lysing the cells and generating cDNA in the compartments, amplifying the barcode templates, attaching barcode sequences to cDNA fragments or fragments generated from cDNA such that a plurality of barcode-attached fragments share the same one or more barcode sequences present in the compartment; removing the compartments and collecting the barcode-attached fragments; and sequencing the barcodes and barcoded nucleic acids to characterize the cDNA profile on a single cell basis. In one embodiment of the method, unique molecular identifier (UMI) sequences can be incorporated into the primers used for cDNA synthesis.
[0047] In another aspect, methods for single cell RNA sequencing are provided, wherein the methods comprise: performing RNA reverse transcription in situ; tagging cDNA in situ; compartmentalizing the processed cells and barcode templates, wherein each compartment contains a processed cell and one or more barcode templates; amplifying the barcode templates and tagged cDNA and coupling the amplified barcode templates to the tagged cDNA in the compartments; removing the compartments and collecting the barcode-attached fragments; and sequencing the barcodes and barcoded nucleic acids to characterize the RNA profile on a single cell basis. In some embodiments, nuclei rather than cells are used as the input material.
[0048] In another aspect, methods for single cell RNA sequencing are provided, wherein the methods comprise: providing a plurality of cells; fixing and / or permeabilizing the cells; providing a reverse transcriptase and providing a plurality of primers, wherein the primers are capable of serving as primers for cDNA synthesis; generating first and second strand cDNA in situ; providing a plurality of transposomes, each transposome comprising at least one transposon and a transposase, tagging double-stranded cDNA in situ; providing a plurality of unique barcode templates; compartmentalizing the processed cells, barcode templates, and primers to generate two or more compartments, the compartments containing cells, one or more barcode templates having different barcode sequences, and primers; in the compartments, amplifying the barcode templates and cDNA fragments, attaching the barcode sequences to the cDNA fragments or fragments generated from the cDNA such that a plurality of barcode-attached fragments share the same one or more barcode sequences present in the compartments; removing the compartments and collecting the barcode-attached fragments; and sequencing the barcodes and barcoded nucleic acids to characterize the cDNA profile on a single cell basis. In some embodiments, nuclei rather than cells are used as the input material. In one embodiment, unique molecular identifier (UMI) sequences can be incorporated into the primers used for cDNA synthesis.
[0049] In one aspect, methods for single-cell RNA sequencing are provided, wherein the methods comprise: providing a plurality of cells, fixing and / or permeabilizing the cells; providing a reverse transcriptase and a plurality of primers, wherein the primers are capable of serving as primers for cDNA synthesis; generating first-strand cDNA in situ; providing a plurality of transposomes, each transposome comprising at least one transposon and a transposase, and tagging RNA / cDNA hybrids in situ; compartmentalizing the cells, barcode templates, and primers to generate two or more compartments, the compartments comprising a cell or nucleus, one or more barcode templates having different barcode sequences, and primers; in the compartments, amplifying the barcode templates and tagged cDNA fragments, attaching the barcode sequences to the cDNA fragments or fragments generated from the cDNA such that a plurality of barcode-attached fragments share the same one or more barcode sequences present in the compartment; removing the compartments and collecting the barcode-attached fragments; and sequencing the barcodes and barcoded nucleic acids to characterize the cDNA profile on a single-cell basis. In some embodiments, nuclei are used instead of cells as the input material. In one embodiment, unique molecular identifier (UMI) sequences can be incorporated into the primers used for cDNA synthesis.
[0050] In one aspect, methods for simultaneously analyzing RNA and DNA in single cells are provided, wherein the methods comprise performing in situ reverse transcription on a plurality of cells before or after cell fixation; performing an in situ strand transfer reaction on the fixed cells; encapsulating the cells individually with one or more barcode templates in compartments; amplifying the barcode templates, cDNA, and DNA fragments in the compartments; coupling the amplified barcode templates to the cDNA and DNA fragments in the compartments; removing the compartments and collecting the barcode-attached fragments; and sequencing the barcodes and barcoded nucleic acids to characterize the RNA and DNA profiles on a single-cell basis. In some embodiments, nuclei are used instead of cells as the input material.
[0051] In one aspect, methods for simultaneously analyzing gene expression and gene regulation in a single cell or for simultaneously analyzing RNA-seq and ATAC-seq in a single cell are provided, wherein the methods comprise: performing in situ reverse transcription on a plurality of cells; performing an in situ strand transfer reaction on the cells; encapsulating the cells individually with one or more barcode templates in compartments; amplifying the barcode templates, cDNA, and accessible chromatin DNA fragments in the compartments; coupling the amplified barcode templates to the cDNA and chromatin DNA fragments in the compartments; removing the compartments and collecting the barcode-attached fragments; and sequencing the barcodes and barcoded nucleic acids to characterize the RNA and accessible chromatin DNA profiles on a single-cell basis. In some embodiments, the in situ strand transfer reaction is performed before the reverse transcription reaction. In some embodiments of the method, the cells are fixed before encapsulation.
[0052] On the other hand, a method for identifying the compartment origin of any barcode when partitioning a barcode template and barcoded targets when there are more than one barcode in a compartment is provided. The method includes providing specific information on the compartment contents, identifying the barcode information of the target and the compartment contents information of the barcode, and grouping barcodes with the same compartment contents information to collect all targets associated with these barcodes.
[0053] In one embodiment, the compartment contents information is the common breakpoint coordinates of tagged fragments from more than one nucleic acid fragment, or the common UMI sequence from more than one target, or a combination thereof.
[0054] The compositions and substances described in the present disclosure and embodiments are isolated or otherwise made in connection with the examples provided herein. Other features and advantages of the present disclosure and embodiments will be apparent from the detailed description and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 A schematic diagram of a nucleic acid barcoding method using a transposome and a barcode template and performing a compartmentalization reaction according to the method is provided. BC represents the barcode on the barcode template.
[0056] Figures 2A - 2D A schematic diagram is provided showing a method of attaching an amplified barcode template to a tagged nucleic acid fragment in a compartment according to the described method. Figure 2A In, the amplified barcode template is used as a primer to further amplify the target (200) to attach the barcode to the target (201) in the compartment. Figure 2B In, the amplified barcode is indirectly coupled to the target (200) using an adapter oligomer (203) such that after amplification, the barcode sequence is attached to the target (202). Figure 2C In, the barcode template and the target (200) are separately double amplified (204, 205) in the compartment, and then the amplified barcode sequence is coupled to the amplified target (206, 207). Figure 2D In, two barcode templates and the target (200) are separately double amplified (210, 213) in the compartment, and subsequently the amplified barcode sequences are coupled to the amplified targets (214, 215). BC represents the barcode on the barcode template. BC1 and BC2 are different barcode sequences.
[0057] Figure 3 A schematic diagram of a single cell ATAC-seq library preparation method is provided. The method includes using a transposome to tag targets in the cell nucleus and coupling the targets to multiple barcode templates through a compartmentalization reaction.
[0058] Figure 4 Provided is a schematic diagram of a single-cell whole-genome barcoding method. The method includes using a transposome to tag targets in the nucleus of fixed cells and coupling the targets to a barcode template through a compartmentalized reaction.
[0059] Figure 5 Provided is a schematic diagram of a method for enriching a targeted region using barcoded nucleic acid fragments and a set of target-specific primers.
[0060] Figure 6 Provided is a schematic diagram showing that single cells barcoded using the methods of the present disclosure can significantly improve the detection ability of somatic mutations. The method involves combining single-cell identification and sequencing error correction with unique molecule identification (UMI). As shown, after sorting using cell IDs, low-frequency mutant genotypes can be identified from mutant cells with minimal background signal in a large number of normal cells. After correction using UMI, the noise caused by sequencing errors is minimized.
[0061] Figure 7 Provided is a schematic diagram illustrating a single-cell RNA-seq method. The method involves in situ reactions and compartmentalized barcode amplification and coupling reactions.
[0062] Figure 8 Provided is a flowchart showing a method for generating a single-cell sequencing library for 5'-end and 3'-end RNA-seq in the same cell.
[0063] Figure 9 Provided is a schematic diagram of a single-cell nucleic acid barcoding reaction for targeted sequencing in a compartment.
[0064] Figure 10 Provided is a flowchart showing the sequencing library preparation workflow for in-cell ATAC-seq and 3'-RNA-seq analysis.
[0065] Figure 11 Provided are graphs of 3'-single-cell RNA-seq analysis using a human and mouse cell mock mixture (1:1 ratio). These graphs support the single-cell characteristics, low collision rate, and scalability of the methods of the present disclosure.
[0066] Figure 12Charts are provided that visualize 3' single-cell RNA-seq data using Uniform Manifold Approximation and Projection (UMAP). These charts support that the methods of the present disclosure can resolve cellular diversity in complex mixtures such as human peripheral blood mononuclear cells (PBMCs). These charts also highlight the enhanced sensitivity of the methods of the present disclosure in identifying rare cell populations.
[0067] Figure 13 The provided charts show profiling of full-length transcriptome information in human Jurkat cells using the methods of the present disclosure.
[0068] Figure 14 The provided schematic diagrams and charts show genome-based bacterial species identification and quantification in five bacterial cell mock mixtures (1:1:1:1:1 ratio) using the methods of the present disclosure. In the chart of the rightmost subfigure, the bars representing bacteria are listed in the following order, from top to bottom: Klebsiella aerogene, Escherichia coli, Citrobacter freundii, Staphylococcus epidermis, Bacillus subtilis.
[0069] The transposase shown in the figure is for illustration only and is shown in the form of a tetramer or dimer. Different transposases can be used in the reaction. Detailed Description
[0070] Improved single-cell nucleic acid detection and sequencing methods are described and characterized herein. The present disclosure is at least partially based on the discovery that amplification of nucleic acid targets during the tagging process offers many advantages compared to conventional sequencing techniques, including enhanced detection of rare genetic variants and allowing full-length sequencing of longer nucleic acid targets using only short-read sequencing techniques.
[0071] Most commercially available sequencing technologies have limited sequencing read lengths. Second-generation high-throughput sequencing technologies can only sequence a few hundred bases and rarely reach thousands of bases. However, the nucleic acid sequences of genes can span from several thousand bases to dozens and hundreds of thousands of bases, which means that sequencing read lengths of dozens of thousands of bases are necessary for successfully determining the haplotypes of all genes.
[0072] Currently, although individual cells are different, most sequencing methods involve extracting DNA or RNA from many cells at once for bulk sequencing. By using averaged molecular or phenotypic measurements of a cell population to represent the behavior of individual cells, conclusions can be affected by the expression profiles of the majority cell population or overexpressed outliers. In addition, such measurements lack the sensitivity to identify all the unique patterns of individual cells, which can reflect the unique functional behavior of cells at a given location and time. Moreover, due to the presence of a high background wild-type signal from normal cells or tissues, the ability to detect very low-frequency somatic mutations is limited, which greatly restricts the ability to use current methods for early tumor detection. However, with the improved ability to identify each single cell provided by the methods described herein, advantageously, we will be able to separate mutant tumor cells from wild-type or normal cells by genotyping at the single-cell level. This approach will result in almost complete elimination of the wild-type background signal generated by normal cells, making somatic mutation detection as easy as germline mutation detection.
[0073] Tn5 transposomes and MuA transposomes have been previously described to simultaneously fragment DNA and introduce adapters at high frequency in vitro to create sequencing libraries for next-generation DNA sequencing (Adey et al. 2010, Caruccio et al. 2011, and Kavanagh et al. 2013). Due to the fragmentation of DNA, these particular protocols remove any phasing or adjacency information. In these protocols, after the DNA reacts with the transposome, column purification, heat treatment steps, protease treatment, or incubation with SDS solution or EDTA solution is required to release the transposase from the strand transfer complex (STC) such that the DNA is tagged to the fragments. The MuA transposome is known to form very stable STCs when attacking DNA targets (Surette et al. 1987, Mizuuchi et al. 1992, Savilahti et al. 1995, Burton and Baker 2003, Au et al. 2004). A similar stability has also been observed for the Tn5 transposome during the transposition reaction (Amini et al. 2014).
[0074] In some embodiments, the present disclosure takes advantage of the stability of STCs (such as Tn5 transposomes and MuA transposomes) and the generation of clonal barcodes by compartmentalized amplification to provide methods for uniquely barcoding nucleic acid target sub-fragments and / or barcoded nucleic acid targets in single cells. Definitions
[0075] Unless otherwise defined, all technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the art to which this disclosure pertains. The following references provide a general definition of many of the terms used in this disclosure and in the embodiments herein: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). Unless otherwise noted, the following terms used herein have the meanings given below.
[0076] As used herein, the term "adapter" refers to a nucleic acid sequence added to a nucleic acid, e.g., by ligation. An adapter can include a primer binding sequence, a barcode, a linker sequence, a sequence complementary to the linker sequence, a capture sequence, a sequence complementary to the capture sequence, a restriction site, an affinity moiety, a unique molecular identifier, and combinations thereof.
[0077] As used herein, the term "amplification" refers to the process of generating multiple copies of an original template. Methods for amplification can include, for example, PCR, RPA, MALBAC, and processes of isothermal amplification methods for linear and exponential amplification.
[0078] As used herein, the term "barcode template" refers to a barcode sequence flanked on one end side by at least one handle sequence or flanked on both end sides by two handle sequences. The length of the barcode sequence can range from 4 bases to 100 bases. The handle sequence can be used as a binding site for hybridization or annealing, as a priming site during amplification, or as a binding site for a sequencing primer or transposase. The barcode sequence can be selected from a pool of known nucleotide sequences or can be randomly selected from randomly synthesized nucleotide sequences. The barcode template can be DNA, RNA, or a DNA / RNA hybrid.
[0079] "Biological sample" refers to any suitable biological sample, including blood and other liquid samples of biological origin, including but not limited to peripheral blood, serum, plasma, cerebrospinal fluid (CFS), urine, feces, saliva, sputum, tears, lavage fluid, synovial fluid. The sample may include cells, tissues, organs or preparations thereof obtained by procedures known and commonly used in the art. The biological samples, cells and / or cell nuclei of the present disclosure can be obtained without limitation from mammals, non-human mammals, humans or non-mammals.
[0080] "Cell unit" refers to a single cell. According to this definition, a single cell includes a physical cell and a virtual cell. For example, a cell unit can be a single cell, a single cell in a compartment, or a data representation of a single cell.
[0081] "Compartment" refers to
[0082] "De novo sequencing" refers to sequencing a new genome for which there is no reference sequence available for alignment. For example, sequence reads are assembled into contigs, and the coverage quality of de novo sequence data depends on the size and continuity of the contigs (i.e., the number of gaps in the data).
[0083] "Haplotype phasing" or "haplotype estimation" refers to determining haplotypes based on genotype data (e.g., genomic DNA), such as determining maternal and paternal haplotypes.
[0084] "Hybridization" refers to the pairing of complementary polynucleotide sequences (e.g., genes as described herein) or portions thereof to form double-stranded molecules under various stringent conditions. (See, e.g., Wahl, G.M. and S.L. Berger (1987) Methods Enzymol. 152:399; Kimmel, A.R. (1987) Methods Enzymol. 152:507). "Hybridization" refers to hydrogen bonding between complementary nucleobases, which can be Watson-Crick, Hoogsteen or reverse Hoogsteen hydrogen bonding. For example, adenine and thymine are complementary nucleobases that pair by forming hydrogen bonds.
[0085] For example, stringent salt concentrations are typically less than about 750 mM NaCl and 75 mM trisodium citrate, preferably less than about 500 mM NaCl and 50 mM trisodium citrate, more preferably less than about 250 mM NaCl and 25 mM trisodium citrate. Low stringency hybridization can be achieved in the absence of organic solvents (e.g., formamide), while high stringency hybridization can be achieved in the presence of at least about 35% formamide, more preferably at least about 50% formamide. Stringent temperature conditions typically include a temperature of at least about 30°C, more preferably at least about 37°C, and most preferably at least about 42°C. Those skilled in the art are familiar with variations in other parameters, such as hybridization time, detergent concentration (e.g., sodium dodecyl sulfate (SDS)), and the addition or exclusion of carrier DNA. These different conditions can be combined as needed to achieve different levels of stringency. In a preferred embodiment, hybridization will occur at 30°C in 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS. In a more preferred embodiment, hybridization will occur at 37°C in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 μg / ml denatured salmon sperm DNA (ssDNA). In a most preferred embodiment, hybridization will occur at 42°C in 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, and 200 μg / ml ssDNA. Those skilled in the art will readily appreciate useful variations of these conditions.
[0086] For most applications, the stringency of the wash step after hybridization will also vary. Wash stringency conditions can be defined by salt concentration and temperature. As described above, wash stringency can be increased by decreasing the salt concentration or increasing the temperature. For example, a stringent salt concentration for the wash step will preferably be less than about 30 mM NaCl and 3 mM trisodium citrate, and most preferably less than about 15 mM NaCl and 1.5 mM trisodium citrate. Stringent temperature conditions for the wash step generally include a temperature of at least about 25°C, more preferably at least about 42°C, and even more preferably at least about 68°C. In one preferred embodiment, the wash step will occur at 25°C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the wash step will occur at 42°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the wash step will occur at 68°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Other variations of these conditions will be readily apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton and Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.
[0087] A "primer set" refers to a set of oligonucleotides that can be used, for example, in polymerase chain reaction (PCR). In some embodiments, a primer set can consist of at least 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 40, 50, 60, 80, 100, 200, 250, 300, 400, 500, 600 or more primers.
[0088] As used herein, the term "transposase" refers to a component of a functional nucleic acid-protein complex that is capable of transposition and mediates transposition, including but not limited to Tn, Mu, Ty, and Tc transposases. The term "transposase" also refers to integrases of retrotransposon or retroviral origin. Transposases also refer to wild-type proteins, mutant proteins, and tagged fusion proteins such as GST-tag, His-tag, etc. and combinations thereof.
[0089] As used herein, the term "transposon" refers to a nucleic acid segment that is recognized by a transposase or integrase and is an essential component of a functional nucleic acid-protein complex that is capable of transposition. The transposon forms a transposome with the transposase and undergoes a transposition reaction. As used herein, the term "transposon" refers to both wild-type and mutant transposons.
[0090] As used herein, "transposable DNA" refers to a nucleic acid segment that contains at least one transposon unit. Transposable DNA can also contain an affinity moiety, unnatural nucleotides, and other modifications. The sequence in transposable DNA other than the transposon sequence can contain linker sequences.
[0091] As used herein, the term "transposome" refers to a stable nucleic acid and protein complex formed by non-covalent binding of a transposase to a transposon. The transposome can contain polymeric units of the same or different monomeric units.
[0092] As used herein, "transposon joining strand" refers to the strand of double-stranded transposon DNA that is joined to the target nucleic acid at the insertion site by a transposase.
[0093] As used herein, "transposon complementary strand" refers to the complementary strand of the transposon joining strand in double-stranded transposon DNA.
[0094] As used herein, "strand transfer complex (STC)" refers to a nucleic acid-protein complex that contains a transposome and its target nucleic acid into which the transposon is inserted, wherein the 3' end of the transposon joining strand is covalently linked to its target nucleic acid. STC is a very stable form of nucleic acid and protein complex that can resist extreme heat and high salt in vitro (Burton and Baker, 2003).
[0095] As used herein, "strand transfer reaction" refers to a reaction between a nucleic acid and a transposome, wherein a strand transfer complex (STC) is formed.
[0096] As used herein, "tagmentation reaction" refers to a fragmentation reaction in which a transposome inserts into a target nucleic acid through a strand transfer reaction to form a strand transfer complex; then the strand transfer complex is disrupted under certain conditions, such as protease treatment, heat treatment, or a protein denaturant (such as an SDS solution, guanidine hydrochloride, urea, etc. or a combination thereof), such that the target nucleic acid is broken into smaller fragments (e.g., tagged nucleic acid fragments) with the transposon attached to the ends of the target nucleic acid. Generally speaking, tagmentation encompasses the initial steps in nucleic acid library preparation, in which unfragmented nucleic acids (e.g., DNA, cDNA, gDNA) are lysed / disrupted and tagged for analysis.
[0097] As used herein, "reaction vessel" refers to a substance having a continuous open space for containing a liquid. In some embodiments, the reaction vessel is selected from the group consisting of: tubes, wells, plates, wells in a multi-well plate, slides, spots on a slide, droplets, tubes, channels, bottles, chambers, and flow cells.
[0098] The preparation of a sequencing library may involve an amplification step. Amplification may involve thermal cycling or isothermal amplification (e.g., by RPA or LAMP methods). Crosslinking may involve overlap extension PCR or using ligases to associate multiple amplification products with each other. Amplification can refer to any method using primers and a polymerase capable of replicating a target sequence with reasonable fidelity. Amplification can be carried out by a native or recombinant DNA polymerase, such as TaqGold TM , T7 DNA polymerase, the Klenow fragment of Escherichia coli DNA polymerase, and reverse transcriptase. The preferred amplification method is PCR. The ranges provided herein should be understood as shorthand for all values within that range. For example, a range of 1 to 50 is understood to include any number, combination of numbers, or sub-range selected from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50.
[0099] Encapsulate nucleic acids with a chain transfer complex and a barcode template in water-in-oil emulsion droplets
[0100] The present disclosure provides methods for encapsulating a nucleic acid target and a barcode template in the form of a strand transfer complex (STC) in water-in-oil emulsion droplets to generate barcoded nucleic acid fragments.
[0101] The nucleic acid target reacts with the transposome (101) to a stable strand transfer complex (102) while maintaining the adjacency of the nucleic acid target ( Figure 1)。The nucleic acid target can be double-stranded. In some embodiments, the nucleic acid target is double-stranded DNA. In some embodiments, the nucleic acid target is a DNA / RNA hybrid. The strand transfer reaction may involve multiple nucleic acid targets in a reaction vessel. In some embodiments, one type of transposome (e.g., Tn5 or MuA) is used; in other embodiments, more than one type of transposome (e.g., Tn5 and MuA) is used simultaneously or sequentially. In one embodiment, the nucleic acid target with STC (102) is mixed with multiple barcode templates (103) in solution. In some embodiments, each barcode template has a unique barcode sequence that is different from the barcode sequence in another barcode template. In some embodiments, there are multiple groups of barcode templates, each group having a unique barcode sequence that is different from the barcode sequences of the other templates in the group, where each group contains at least one barcode template. In some embodiments, the barcode templates are oligonucleotides that are freely present in solution. In some embodiments, the barcode templates are arranged in the form of nanospheres. In some embodiments, the barcode templates are encapsulated in droplets. In some embodiments, the barcode templates are immobilized on a carrier, which can be a solid bead or particle (e.g., nanoparticle), or a soluble bead or particle, or a combination thereof. In some embodiments, the carrier contains only a single barcode template. In some embodiments, the carrier contains multiple barcode templates, where each template has a unique barcode sequence that is different from the barcode sequences of the other barcode templates. In some embodiments, the carrier contains only a single group of barcode templates, where the group of barcode templates has the same barcode sequence. In some embodiments, the carrier contains multiple groups of barcode templates, each group having a unique barcode sequence that is different from the barcode sequences of the other templates in the group, where each group includes at least one barcode template.
[0102] At least one transposable DNA in the transposome can hybridize directly ( Figure 2A ) or indirectly through an adaptor and / or primer ( Figure 2B ) to one end of the barcode template. Other enzymes and substrates, such as DNA polymerase, dNTPs, and primers, are also provided in the same reaction vessel in the form of an aqueous solution. In some embodiments, the primers are used to amplify the barcode templates. In some embodiments, the primers can be used to amplify the tagged nucleic acid target fragments. Amplification includes exponential amplification and linear amplification. In some embodiments, different primers can be used to amplify the barcode templates and the tagged nucleic acid target fragments in parallel ( Figure 2C ); thereafter, the two sets of amplification products can be ligated through the homology shared between two internal primers ( Figure 2C, 208 and 209) or incorporated / coupled into one entity through additional linkers capable of bridging the barcode template and the tagged fragments. The water-in-oil emulsion droplets (104) are generated under conditions where one to several nucleic acid targets with STC are mixed with a barcode template in one droplet. Appropriate titration of the nucleic acid targets can be performed using STC and the barcode template based on the Poisson distribution. In some embodiments, multiple barcode templates with different barcode sequences can be used in the emulsion droplets to significantly increase the presence of barcodes in the emulsion droplets, which advantageously increases the number of droplets containing positive products, thus significantly improving the reaction yield of the barcoding reaction.
[0103] In some embodiments, when amplifying the barcode template and the tagged fragments before attaching the barcode sequence to the tagged fragments, if different barcodes are randomly attached to the amplified copies of the tagged fragments ( Figure 2D ), then multiple barcode templates with different barcode sequences used in the same emulsion droplet will not affect the true representation of the nucleic acid targets. Thus, when the barcode template and the nucleic acid targets are encapsulated in the same droplet, most emulsion droplets will contain the barcode template for barcoding the nucleic acid targets. This makes it possible to obtain almost 100% of the droplets containing any nucleic acid target available for barcoding. In an embodiment, the diameter of the emulsion droplets is from 1 μm to 200 μm, or from 5 μm to 30 μm. When multiple barcodes are present in the emulsion droplet compartment, the break point coordinates of the tagged fragments can be used to trace these barcodes back to an original compartment. Specifically, the break points generated by transposase tagging are different between different nucleic acid targets. If a DNA fragment with a barcode shares the same break point coordinates with a fragment with one or more other barcodes, then these fragments may be derived from the same original compartment. For multiple nucleic acid targets in an experiment, it is possible for two different nucleic acid fragments to generate the same break point after transposase tagging. When multiple break points are used for discrimination, the chance of such a conflict is much lower. In some embodiments, uniquely molecular identifier (UMI)-tagged transposomes can be used during the strand transfer reaction or the tagging reaction to increase the uniqueness of the fragments for identification. When different barcodes share many fragments with the same set of UMI populations in addition to the same set of fragment break points, UMI information can be used for compartment identification.
[0104] The STC is processed to release the transposase from the tagged nucleic acid target fragments, for example, by heat treatment. After heat treatment (e.g., at 60 °C to 75 °C for about 5 - 10 minutes), the transposase will be released from the STC and the nucleic acid targets will break into smaller fragments. In some embodiments, when still in the emulsion droplets, DNA polymerase fills in the gaps left during the transposition reaction. Emulsion amplification is performed to amplify the barcode template in the droplets. The amplified barcode template will be directly (Figure 2A ) or indirectly Figure 2B ) hybridizes with the tagged fragment and attaches the barcode sequence to the fragment (105, 201, and 202) during the amplification reaction. In some embodiments, during the emulsion reaction, unique molecular identifiers (UMIs) are added to the barcode template. In some embodiments, the UMIs are incorporated as the adapters (203) or primers (209 and 212) in FIG. 2. After the emulsion amplification reaction, the emulsion droplets are dispersed, for example, with high salt, detergent, alcohol, organic chemicals, or a combination thereof. After the emulsion droplets are dispersed, the aqueous phase of the resulting solution is collected. In some embodiments, one or more biotinylated primers are used such that the amplified barcoded fragments can be readily bound to streptavidin beads. In some embodiments, one or more biotinylated dNTPs are used for emulsion amplification. In some embodiments, during emulsion amplification, primers with sample-specific barcodes are used for the emulsion droplets such that the emulsion amplification products from different sample reactions can be pooled together for final amplification or adapter modification to create a sequencing read library.
[0105] In some embodiments, the nucleic acid target is genomic DNA. This barcoding method can be used for de novo sequencing, genome-wide haplotype phasing, and structural variant detection. In some embodiments, the nucleic acid target is a DNA fragment, cDNA, or a portion of DNA captured by hybridization capture, primer extension, or PCR amplification. This barcoding method is capable of phasing variants of these DNA molecules. In some embodiments, target-specific primers can be used in compartments to amplify specific nucleic acid targets that react or do not react with the transposome.
[0106] Encapsulate transposase-tagged cells or cell nuclei and a barcode template in water-in-oil emulsion droplets
[0107] A method is described herein for encapsulating cells or cell nuclei and barcode templates in water-in-oil emulsion droplets after a strand transfer reaction, and further for generating barcoded nucleic acid fragments for single-cell level analysis.
[0108] Assaying for Transposase-Accessible Chromatin using sequencing (ATAC-seq) has gained popularity as a cutting-edge molecular biology tool for assessing genome-wide chromatin accessibility (Buenrostro et al., 2013). ATAC-seq identifies accessible chromatin regions by tagging open chromatin with a hyperactive mutant Tn5 transposase that integrates sequencing adapters into open regions of the genome. The tagged DNA fragments are purified, amplified by PCR, and sequenced. The sequencing reads are then used to infer regions of increased accessibility and to map regions of transcription factor binding sites and nucleosome positions. Although the activity level of the native wild-type transposase is low, ATAC-seq employs a mutant hyperactive transposase (Reznikoff et al., 2008) that has been successfully adapted for efficient identification of open chromatin and identification of regulatory elements across the genome. In addition, single-cell ATAC-seq is used to dissociate single nuclei and perform the ATAC-seq reaction individually (Buenrostro et al., 2015). Higher-throughput single-cell ATAC-seq uses combinatorial cell indexing to measure chromatin accessibility in thousands of individual cells. Single-cell ATAC-seq can identify cell types and states for developmental lineage tracing. ATAC-seq is likely to be a key component of an integrated epigenomics workflow.
[0109] In some embodiments, the present disclosure includes methods of encapsulating transposase-treated cell nuclei and unique barcode templates using water-in-oil droplet emulsions. The method also includes clonal amplification of the barcode templates within the emulsion droplets and attaching the clonally amplified barcodes to the tagged accessible DNA fragments ( Figure 3 ). The tagged DNA can also be amplified in the emulsion droplets. This barcoding method of the present disclosure provides the advantages of high-throughput and low-cost cell indexing for single-cell ATAC-seq analysis.
[0110] In some embodiments, cell nuclei (302) are collected from a cell or tissue sample (301) and incubated with a transposome (303) to form STC (304), which is then mixed with a plurality of barcode templates (305) in a bulk reaction ( Figure 3)。In some embodiments, each barcode template has a unique barcode sequence that is different from the barcode sequences in other barcode templates. In some embodiments, there are multiple groups of barcode templates, each group having a unique barcode sequence that is different from the barcode sequences of other groups of barcode templates, where each group contains at least one barcode template. In some embodiments, the barcode templates are oligonucleotides that are freely present in solution. In some embodiments, the barcode templates are arranged in the form of nanospheres. In some embodiments, the barcode templates are encapsulated in droplets. In some embodiments, the barcode templates are immobilized on a carrier, which can be a solid bead or particle (e.g., nanoparticle), or a soluble bead or particle, or a combination thereof. In some embodiments, the carrier contains only a single barcode template. In some embodiments, the carrier contains multiple barcode templates, where each template has a unique barcode sequence that is different from the barcode sequences of each other barcode template. In some embodiments, the carrier contains only a single group of barcode templates, where the group of barcode templates has the same barcode sequence. In some embodiments, the carrier contains multiple groups of barcode templates, each group having a unique barcode sequence that is different from the barcode sequences of other groups of barcode templates, where each group contains at least one barcode template.
[0111] In some embodiments, whole cells are treated with a transposome to form STCs in the cell nucleus without separating the cell nucleus. In some embodiments, the transposome contains a mutant hyperactive Tn5 transposase. In some embodiments, the transposome contains MuA transposase. Other enzymes and substrates, such as DNA polymerase, dNTP, and primers (306), are also provided in the same bulk reaction in the form of an aqueous solution. An oil-in-water emulsion droplet (307) is generated under the condition that there is one cell nucleus and one barcode template in most droplets by limiting the distribution or titration based on the Poisson distribution. In an embodiment, the diameter of the emulsion droplet is 10 μm to 200 μm, or 20 μm to 60 μm.
[0112] The STC is processed to release the transposase from the tagged nucleic acid target fragments, e.g., by heat treatment. After heat treatment, e.g., treating at 60 °C to 75 °C for about 5 - 10 minutes, the transposase will be released from the STC and the nucleic acid targets will break into smaller fragments. In some embodiments, when still in the emulsion droplets, the DNA polymerase present in the droplets fills in the gaps left during the transposition reaction. The nuclear membrane is broken during the emulsion PCR denaturation step and emulsion amplification is performed to amplify the barcode templates in the droplets. The amplified barcode templates are capable of hybridizing directly or indirectly to the tagged fragments and attaching the barcode sequences to the fragments during the amplification reaction (308). In some embodiments, the barcoded template and the tagged fragments are first amplified in parallel and then combined or coupled together to form barcoded tagged fragments, as Figure 2C and 2D shown. After the emulsion amplification reaction, the emulsion droplets are dispersed, e.g., by high salt, detergent, alcohol, organic solution, or a combination thereof. After the emulsion droplets are dispersed, the aqueous phase of the resulting solution is collected. In some embodiments, one or more biotinylated primers or one or more biotinylated dNTPs are used, enabling the amplified barcoded fragments to be easily bound to streptavidin beads. In one embodiment, the sequencing library prepared from these barcoded fragments is a single cell ATAC-seq library.
[0113] The present disclosure also provides a single cell whole genome sequencing method as described herein. The method employs emulsion encapsulation of transposase-treated alcohol-fixed cell nuclei and unique barcode templates. The method further includes clonally amplifying the barcode templates within the emulsion droplets and attaching the barcodes to the tagged genomic DNA fragments from the fixed cell nuclei ( Figure 4 ).
[0114] In some embodiments, nuclei (402) are collected from a cell or biological sample (e.g., tissue sample (401)) and fixed. Fixatives such as alcohol-based fixatives or Hepes-glutamate buffer-mediated organic solvent protection effect (HOPE) fixatives, or other similar fixatives, can be used in these methods to stabilize / denature the proteins in the nuclei while maintaining the integrity of the nucleic acid contents in the nuclei (403). In some embodiments, fixation exposes all genomic DNA from chromatin in the nuclei. In some embodiments, the fixed cells are used directly without separating the nuclei. After washing away the fixation solution, the nuclei are treated with a transposome (404) to form STC (405) with genomic DNA, and then mixed with a plurality of different barcode templates (406) in a bulk reaction. Other enzymes and substrates such as DNA polymerase, dNTP, and primers (407) are also provided in the same bulk reaction in the form of an aqueous solution. By limiting the Poisson distribution-based partitioning or titration, water-in-oil emulsion droplets (408) are generated under the condition that there is one nucleus and one barcode template in the droplet. In embodiments, the diameter of the emulsion droplets is from 10 μm to 200 μm, or from 20 μm to 60 μm.
[0115] The STC is treated to release the transposase from the tagged nucleic acid target fragments, for example, by heat treatment. After heat treatment, for example, treating at 60 °C to 75 °C for about 5 - 10 minutes, the transposase will be released from the STC and the nucleic acid targets will break into smaller fragments. In some embodiments, when still in the emulsion droplets, the DNA polymerase present in the droplets fills the gaps left during the transposition reaction. The nuclear membrane of the nuclei is broken, and emulsion amplification is performed to amplify the barcode templates in the droplets. The amplified barcode templates are capable of hybridizing directly or indirectly to the tagged fragments and attaching the barcode sequences to the fragments during the amplification reaction (409). In some embodiments, first the barcode templates and the tagged fragments are amplified in parallel and then combined together to form barcoded tagged fragments, as Figure 2C and 2D . After the emulsion amplification reaction, the emulsion droplets are dispersed, for example, with high salt, detergent, alcohol, organic reagent, or a combination thereof. After the emulsion droplets are dispersed, the aqueous phase of the resulting solution is collected. In some embodiments, one or more biotinylated primers or one or more biotinylated dNTPs are used so that the amplified barcoded fragments can be easily pulled out with streptavidin beads. In some embodiments, the library prepared from these barcoded fragments can be directly used for single-cell whole-genome sequencing and single-cell copy number variation (CNV) analysis. In some embodiments, the library prepared from these barcoded fragments can be used for further targeted capture of the entire exome or targeted capture of smaller targeted regions for targeted sequencing ( Figure 5)。In some embodiments, cells from a metagenomic sample are directly used in the barcoding reaction. In some embodiments, the prokaryotic cell wall can be permeabilized enzymatically and / or chemically. Advantageously, this single-cell sequencing method of the present disclosure eliminates the need for genomic DNA preparation, which is a known bottleneck in metagenomic sample preparation, while directly preserving high-molecular-weight DNA intact in the cells, thereby improving the assembly efficiency. The method of the present disclosure will well preserve the organism composition in the metagenomic sample and utilize barcode-based cell-level information to improve the accuracy of organism composition measurement, rather than only using genomic DNA-level information, which contains more biases due to accessibility, amplification, or sequencing.
[0116] In some embodiments, the cells are microbes. In some embodiments, the cells are microbiome cells or metagenomic cells. In some embodiments, the microbe or metagenomic sample is pretreated with lysozyme or other cell wall lysing enzymes to facilitate cell wall removal as part of the preparation process. In some embodiments, the method is used to analyze metagenomic or microbiome samples for species identification, composition analysis, and association analysis of microbial hosts and their plasmids or phages or viruses.
[0117] One advantage of the single-cell targeted barcoding and / or sequencing methods disclosed herein is that, compared to known barcoding and / or sequencing methods, they have much higher sensitivity for detecting low-frequency genetic variations (e.g., detecting somatic mutations) ( Figure 6 ). Since this method can uniquely barcode individual cells and can detect any mutation at the single-cell level, this will effectively eliminate the background noise from surrounding cells. This provides extremely high sensitivity for detecting extremely low-frequency somatic mutations, which is exactly what is needed for early cancer detection. Figure 6 Illustrates the advantages of genotyping at the single-cell level. Figure 6 Shows that the cell contains a mutant allele A (601), but in the same sample, there are many wild-type cells containing the normal allele T (602). During the tagging reaction, unique molecular identifiers (UMIs) are added. By incorporating molecule-specific UMIs during single-cell barcoding and sequencing, the sequencing reads can first be grouped based on their cell ID, and for each cell, sequencing errors can be identified based on the UMIs and correct variant identification can be easily performed. This method can be applied to circulating tumor cells, tissue biopsy samples, or tissue sections.
[0118] In some embodiments, multiple barcode templates with different barcode sequences may be present in emulsion droplets to increase the capture rate. When multiple barcodes are present in an emulsion droplet and are shared by a single nucleus or cell, these barcodes can be traced back to their original nucleus or cell by leveraging the breakpoint coordinates of the tagged fragments. Specifically, the breakpoints generated by transposase tagging are different between different nuclei or cells. If a DNA fragment attached with a barcode shares the same breakpoint coordinates with a fragment attached with one or more other barcodes, these barcodes may originate from the same original nucleus or cell. After transposase tagging, two nuclei or cells may generate the same breakpoint in some fragments. When multiple breakpoints are used for discrimination, the chance of such a conflict is much lower. The more breakpoint coordinates shared between two barcodes, the higher the confidence that these two barcodes come from the same compartment (i.e., the same cell or nucleus). In some embodiments, the randomness of the tagged breakpoints is used as a UMI function to track duplicates caused by amplification and improve the counting accuracy of unique targets.
[0119] When multiple different barcode templates are present in droplets that capture cells or nuclei, subsequent amplification that combines the barcode templates and the tagged cell contents creates additional copies of the same cell contents, which can be DNA, RNA, or other cellular targets, and these additional copies are randomly conjugated to different barcode templates. When multiple copies of the cell contents are randomly shared and captured by different barcode templates in a single droplet, such that each barcode template (or group of templates) can capture sufficient cell contents to represent the cell or nucleus in the droplet, this can effectively amplify the signal from a single cell, thereby creating an "amplified" copy of the single cell. Although there is only one cell or nucleus in the droplet, multiple cells or nuclei representing the same cell or nucleus are created after amplification in the droplet by the multiple barcode templates in the droplet. This single-cell amplification can improve downstream clustering analysis for cell population characterization and increase the assay sensitivity for detecting rare cell populations with a low number of input cells or nuclei in single-cell responses. The methods of the present disclosure provide new single-cell library methods capable of amplifying single cells for use in the art.
[0120] In some embodiments, the methods of the present disclosure can also be used for single-cell RNA analysis. In some embodiments, a reverse transcriptase and cDNA primers as a first set of primers can be included in an emulsion reaction. In some embodiments, the cDNA primers contain a poly-T sequence at the 3' end; in some embodiments, the cDNA primers have a GGG nucleotide sequence at the 3' end; in some embodiments, the cDNA primers have a target-specific primer at the 3' end. In some embodiments, cDNA is synthesized using mRNA as a template; in some embodiments, cDNA is synthesized using other RNA species as a template. In the early stage of the emulsion reaction, the reverse transcriptase generates cDNA or partial cDNA from mRNA in a single cell or nucleus. Then, barcode tagging is performed in the same manner as any of the previously described methods, except that cDNA is used as the input DNA. By using different primers for reverse transcription or cDNA priming, this method can be used for single-cell transcriptome analysis, single-cell 3'-RNA-Seq analysis, single-cell 5'-RNA-Seq analysis, single-cell target-sequencing (target-seq) applications, and immune repertoire analysis. The methods of the present disclosure can combine in situ reactions of a large number of cells and compartmentalized amplification and barcode tagging reactions by encapsulating individual processed cells with one or more barcode templates, thereby enabling high-throughput single-cell RNA analysis.
[0121] Figure 7Shows an embodiment of a single-cell RNA barcoding method according to the present disclosure. A cell (701) is first permeabilized (702). In some embodiments, the RNA in the permeabilized cell (702) is transcribed into cDNA in situ (703) by reverse transcriptase. A second strand of DNA is synthesized to form double-stranded DNA, which serves as the input for in situ tagging. In some embodiments, the RNA in the cell is transcribed into first-strand cDNA in situ by reverse transcriptase. The RNA / cDNA hybrid duplex can be used as the input for in situ tagging (704). In some embodiments, the cDNA primer has a poly-T sequence at the 3' end; in some embodiments, the cDNA primer has a GGG sequence at the 3' end; in some embodiments, the cDNA primer has a target-specific primer at the 3' end; in some embodiments, mRNA is used as a template to synthesize cDNA; in some embodiments, other RNA species are used as templates to synthesize cDNA. The processed cells containing the in situ-tagged cDNA (704) are encapsulated with one or more barcode templates (705) for a cloning amplification reaction. During the cloning reaction, the tagged cDNA fragments (706) are released from the cells, both the one or more barcode templates and the tagged cDNA are amplified (dual amplification), and the amplified barcode templates (707) are coupled to the amplified cDNA fragments (708) and generate multiple barcode-attached fragments that share the same one or more barcode sequences present in the compartment (709). By using different primers for reverse transcription or cDNA priming, this method can be used for single-cell transcriptome analysis, single-cell 3' RNA-Seq analysis, single-cell 5' RNA-Seq analysis, single-cell targeted-sequencing (target-seq) applications, and immune repertoire analysis.
[0122] In some embodiments, 3' end RNA and 5' end RNA targets can be captured for the same cell in the same experiment, i.e., 3' RNA-seq and 5' RNA-seq analyses are performed simultaneously ( Figure 8)。In some embodiments, tagged cDNA / RNA hybrids or double-stranded cDNA can be used to capture full-length transcripts for full-length single-cell transcriptome analysis. Full-length transcripts and / or transcriptome analysis are very useful for studying alternative splicing and mRNA isoforms. In some embodiments, optimized fixation and / or permeabilization conditions are designed to drive the reaction to cytoplasmic mRNA, mainly for full-length transcriptome analysis, to reduce the representation of precursor mRNA and genomic or chromatin DNA in the nucleus. The single-cell full-length transcriptome method described herein is particularly suitable for short-read sequencing platforms because the method can break long full-length transcripts into multiple short fragments and couple them to the same barcode template in a cell unit. Short transcript fragments are easy to amplify, especially compared to longer sequences, and the final library length can be relatively short. The resulting library is very suitable for short-read sequencing platforms. In some embodiments, when the tagged transcripts are kept long, the method can be used for long-read sequencing platforms.
[0123] In some embodiments, multiple barcode templates with different barcode sequences can be present in the emulsion droplets to increase the cell capture rate. When multiple barcode templates are present in the emulsion droplets and are shared by one cell or nucleus in a compartment, these barcodes can be traced back to an original cell / nucleus through UMIs on the reverse transcription primers or through unique tagged breakpoints on the transcripts. In some embodiments, it is preferred to keep these barcodes as (virtual) separate cells rather than merging these different barcodes back to their original cell source. In this case, one cell or nucleus may expand into multiple cells or nuclei after the reaction. These expanded cells can improve downstream clustering analysis for cell population characterization and increase detection sensitivity when detecting rare cell populations with a small number of input cells or nuclei in single-cell reactions.
[0124] Encapsulate cells, a barcode template, and target-specific primers in water-in-oil emulsion droplets
[0125] The present disclosure also provides a high-throughput method for single-cell targeted sequencing. Figure 9 An embodiment of the high-throughput method is illustrated. Isolated cells or nuclei (902) are encapsulated in emulsion droplets together with unique barcode templates (903) and a first set of target-specific primers (904) Figure 9,901). Other enzymes and substrates, such as DNA polymerase, dNTP, and common primer, can also be provided in the form of an aqueous solution. The water-in-oil emulsion droplets (901) are generated under the condition that there is one cell or one nucleus and one barcode template in the droplets through limited titration or partitioning based on Poisson distribution. In an embodiment, the diameter of the emulsion droplets is from 10 μm to 200 μm, or from 20 μm to 100 μm. The cell membrane and / or nuclear membrane are broken to release genomic DNA into the emulsion droplets. An emulsion amplification reaction is carried out to amplify the barcode template and attach target-specific primers to the barcode template in the droplets. The single-stranded amplified barcode template (905) having a target-specific sequence at the 3' end is capable of hybridizing to the genomic DNA target and making copies of the targeted region during the emulsion amplification reaction. In some embodiments, during the generation of the emulsion droplets, a second set of target-specific primers (906) is included in the aqueous solution. After the emulsion amplification reaction, barcoded amplicons (907) of the target will be generated, which can be used for sequencing library preparation and sequencing analysis. In some embodiments, to reduce primer dimers generated during amplification, primers containing dUTP can be used and combined with treatment with UDG / APE1 / ExoI after emulsion amplification. After cleaning up the primer dimers, sequencing library adapters can be added by ligation.
[0126] Method for performing RNA and DNA analysis in the same cell
[0127] Currently, most single-cell analysis methods can only perform RNA or DNA analysis on different single cells separately. In other words, currently known single-cell analysis methods cannot analyze RNA and DNA from the same cell simultaneously.
[0128] However, the methods of the present disclosure include simultaneously monitoring RNA expression and determining DNA genotype of the same cell. In some embodiments, the cells are fixed after in situ reverse transcription reaction to generate cDNA to dissociate DNA and / or stabilized products from proteins. In some embodiments, the cells are first fixed before performing the in situ reverse transcription reaction. Poly-T primers can be used to capture 3’mRNA. In some embodiments, UMI sequences are associated with the poly-T primers. Either a strand transfer reaction or a tagging reaction can be performed in situ within the treated cells or after the cells are encapsulated in compartments with barcoded templates. In some embodiments, if the nucleic acid targets are all specific, neither a strand transfer reaction nor a tagging reaction is required. During cell encapsulation within the compartments, cDNA-specific primers and DNA target-specific primers and / or transposon-specific primers are included simultaneously with the primers for amplifying the barcoded templates. In some embodiments, when poly-T primers are used, cDNA amplification is directed to 3’mRNA. In some embodiments, DNA amplification is target-specific or genome-wide specific. After amplification of one or more barcoded templates and cDNA and / or DNA fragments, the barcoded templates are ligated to the amplified cDNA and / or DNA fragments in the compartments. Subsequently, the barcoded cDNA and DNA are released from the compartments and collected for further analysis of gene expression and genomic variation.
[0129] The present disclosure also provides a method for simultaneously performing ATAC-seq and RNA-seq on the same cell. The cells are permeabilized and reverse transcribed in situ using poly-T labeled primers to generate cDNA. In some embodiments, cDNA is generated only after the first strand cDNA. In some embodiments, the cDNA is generated after the second strand cDNA synthesis. These cells are incubated with a transposome to perform a strand transfer reaction at the open chromatin sites in the cell nucleus and the cDNA in the cells. In some embodiments, the strand transfer reaction at the open chromatin sites is performed before reverse transcription. Then, the cells are individually encapsulated in compartments with one or more barcoded templates for barcoded amplification and tagged RNA and DNA amplification. In some embodiments, before encapsulation, these cells are fixed to denature cellular proteins and exogenous reverse transcriptase and transposase. In some embodiments, the cell nuclei are isolated from the cells before the strand transfer reaction and / or reverse transcription reaction ( Figure 10 ).
[0130] Transcriptome and epitope cellular indexing by sequencing (CITE-seq) is a multimodal single-cell phenotyping method that uses DNA-barcoded antibodies to convert the detection of proteins into quantitative, sequenceable readouts. The oligonucleotides conjugated to the antibodies serve as synthetic transcripts and are captured in most large-scale oligonucleotide-dT-based single-cell RNA-seq library preparation protocols (Stoeckius et al., 2017). In some embodiments, CITE-seq libraries can be efficiently generated when cDNA primers are labeled with a polyT sequence.
[0131] In some embodiments, the encapsulated target is not a nucleic acid, genome, protein, nucleus, cell, or microbe, but a protein complex, protein and nucleic acid complex, small molecule, macromolecule, compound, ligand, particle, microparticle, or a combination thereof. The encapsulated target can be labeled with a nucleic acid or attached to a nucleic acid as an identifiable label or marker.
[0132] In some embodiments of the methods of the present disclosure, the cell is a eukaryotic cell; in other embodiments, the cell is a prokaryotic cell.
[0133] Encapsulation in a water-in-oil emulsion is a compartmentalization (isolation) method used in the methods of the present disclosure, but other isolation methods are also feasible and can be used in the methods. Certain types of liposomes, for example, giant unilamellar vesicles (GUVs) with a diameter of 1-200 μm, have shown very high thermal stability and are capable of performing PCR amplification within their outer shell (Kurihara et al., 2011, Laouini et al., 2012). Thus, in some embodiments, GUVs can be used as compartments in the present method. In some embodiments, compartmentalization is achieved through micropores. In some embodiments, compartmentalization is achieved through an open array. In some embodiments, compartmentalization is achieved through a microarray, microtiter plate, or other physically separated compartmentalization method.
[0134] One embodiment relates to a method for analyzing and / or counting nucleic acids from single cells, the method comprising (a) providing a sample comprising cells within a plurality of cells, wherein the cells comprise a plurality of sample nucleic acids; (b) generating a plurality of barcoded polynucleotides from the plurality of sample nucleic acids of the cells, wherein the barcoded polynucleotides comprise a barcode sequence configured to distinguish the sample nucleic acids from other sample nucleic acids in other cells; and a sample sequence from the sample nucleic acids in the cells, wherein the sample sequence comprises a sequence distinguishable from other sample sequences of other sample nucleic acids in the cells; (c) sequencing the barcoded polynucleotides to determine the sample sequence and the barcode sequence; (d) analyzing and / or counting the sample nucleic acids in the cells with the barcode sequence and sample sequence information. In some embodiments, the method further comprises generating a plurality of compartments, wherein the cells are individually compartmentalized in the compartments before or during step (b). In some embodiments, the method further comprises amplifying the barcoded polynucleotides to generate a plurality of amplified barcoded polynucleotides before step (c). In some embodiments, the compartment forms include droplets, emulsion droplets, liposomes, microwells, pores, microarrays, open arrays, microtiter plates, or combinations thereof. In some embodiments, the sample nucleic acids are selected from total DNA, a portion of DNA, total RNA, a portion of RNA, and combinations thereof in the cells. In some embodiments, the plurality of barcoded polynucleotides are generated by reactions selected from ligation, hybridization, strand transfer reactions, transposition, tagging, primer extension, reverse transcription, amplification, and combinations thereof. In some embodiments, the sample nucleic acids in the cells are pretreated in situ before step (b) for reverse transcription, transposition, tagging, strand transfer reaction, ligation, hybridization, restriction endonuclease digestion, crosslinking, fixation, or combinations thereof. In some embodiments, the sample sequences with distinguishable sequences are generated by strand transfer, transposition, tagging, random priming, random reverse transcription, random digestion, or combinations thereof. In some embodiments, the sample sequences with distinguishable sequences are used as unique molecular identifiers for the sample nucleic acids. In some embodiments, at least 80% of the sample sequences with distinguishable sequences comprise unique sequences different from other sample sequences in the cells. In some embodiments, at least 90% of the sample sequences with distinguishable sequences comprise unique sequences different from other sample sequences in the cells. In some embodiments, step (d) further comprises using the barcode sequence to identify the cellular origin of the sample nucleic acids and using the sample sequence to determine the uniqueness of the sample nucleic acids from other sample nucleic acids in the cells. In some embodiments, the cells consist essentially of cell nuclei isolated from the cells.
[0135] One embodiment relates to a method for generating barcoded polynucleotides from cell-based DNA or RNA, comprising: (a) providing a sample comprising a plurality of cells, wherein the cells comprise a plurality of sample DNA or sample RNA; (b) generating a plurality of first barcoded polynucleotides from the plurality of sample DNA of the cells, and generating a plurality of second barcoded polynucleotides from the plurality of sample RNA of the cells, wherein the first barcoded polynucleotides from the sample DNA comprise: a sample sequence from the sample DNA in the cell; a barcode sequence for distinguishing the sample DNA from other sample DNA in different cells; and a sample DNA-specific adapter sequence, wherein the adapter sequence comprises the same first barcoded polynucleotide from the sample DNA; wherein the second barcoded polynucleotides from the sample RNA comprise a sample sequence from the sample RNA in the cell; a barcode sequence for distinguishing the sample RNA from other sample RNA in different cells; a sample RNA-specific adapter sequence, wherein the adapter sequence comprises the same second barcoded polynucleotide from the sample RNA; (c) sequencing the first and second barcoded polynucleotides to determine the sample sequence and the barcode sequence; (d) analyzing the sample DNA and sample RNA in the cells using the barcode sequence and sample sequence information. In some embodiments, the method further comprises generating a plurality of compartments, wherein the cells are individually partitioned into the compartments before or during step (b). In some embodiments, the method further comprises amplifying the first and second barcoded polynucleotides before step (c) to generate a plurality of amplified first and second barcoded polynucleotides. In some embodiments, the compartment form comprises droplets, emulsion droplets, liposomes, microwells, pores, microarrays, open arrays, microtiter plates, or combinations thereof. In some embodiments, the sample DNA is the total DNA of the cell, a portion of the DNA, or accessible chromatin DNA. In some embodiments, the sample RNA is the total RNA of the cell, a portion of the RNA, or mRNA. In some embodiments, the plurality of first and second barcoded polynucleotides are generated by a reaction selected from the group consisting of ligation, hybridization, strand transfer reaction, transposition, tagging, primer extension, reverse transcription, amplification, and combinations thereof. In some embodiments, the sample DNA in the cells is pre-treated in situ before step (b) for strand transfer reaction, transposition, tagging, ligation, hybridization, restriction enzyme digestion, cross-linking, fixation, or combinations thereof. In some embodiments, the sample RNA in the cells is pre-treated in situ before step (b) for reverse transcription, strand transfer reaction, transposition, tagging, ligation, hybridization, restriction endonuclease digestion, cross-linking, fixation, or combinations thereof. In some embodiments, the sample sequence from the first barcoded polynucleotide is a sequence distinguishable from other sample sequences of other sample DNA in the cell.In some embodiments, the sample sequence from the second coded polynucleotide is a sequence distinguishable from other sample sequences of other sample RNAs in the cell. In some embodiments, the sample sequence with a distinguishable sequence is generated by a strand transfer reaction, transposition, tagging, random priming, random reverse transcription, random digestion, or a combination thereof. In some embodiments, the sample sequence with a distinguishable sequence serves as a unique molecular identifier of the sample DNA or sample RNA. In some embodiments, at least 80% of the sample sequences with a distinguishable sequence contain a unique sequence different from other sample sequences in the cell. In some embodiments, at least 90% of the sample sequences with a distinguishable sequence contain a unique sequence different from other sample sequences in the cell. In some embodiments, the barcode sequences between the first and second coded polynucleotides in the cell are the same. In some embodiments, step (d) further comprises using the barcode sequence to identify the common cellular origin of the sample DNA or sample RNA, and using the sample sequence to characterize the sample DNA and the sample RNA in the cell. In some embodiments, the cell consists essentially of a cell nucleus isolated from the cell.
[0136] One embodiment relates to a method for tracking the origin of targets by adding barcode labels, comprising: (a) isolating one or more unique barcode templates from targets in a compartment; (b) amplifying the barcode templates and modifying the targets, wherein the modified targets are configured to ligate to the barcode templates in the compartment; (c) generating barcode-labeled modified targets, wherein multiple modified targets share the same one or more barcode sequences present in the compartment; and (d) removing the separation between compartments and collecting the barcode-labeled modified targets for sequencing characterization. In some embodiments, the method further comprises identifying the compartment origin of different barcode sequences present in the same compartment based on shared compartment contents. In some embodiments, the targets are selected from nucleic acids, proteins, protein complexes, protein and nucleic acid complexes, ligands, compounds, cell nuclei, cells, microorganisms, small molecules, macromolecules, particles, microparticles, and combinations thereof. In some embodiments, the modification of the targets is selected from strand transfer reactions, transposition, tagging, reverse transcription, amplification, primer extension, restriction enzyme digestion, hybridization, ligation, fragmentation, crosslinking, and combinations thereof. In some embodiments, the targets are processed and / or modified prior to encapsulation, wherein the processing is selected from denaturation, permeabilization, fixation, tagging, antibody conjugation, in situ reactions, and combinations thereof; and the modification is selected from strand transfer reactions, transposition, tagging, reverse transcription, amplification, primer extension, restriction enzyme digestion, hybridization, ligation, fragmentation, crosslinking, and combinations thereof. In some embodiments, the isolation compartments are selected from droplets, emulsion droplets, liposomes, microwells, open arrays, microtiter plates, and combinations thereof. In some embodiments, the barcode template comprises a barcode sequence and at least one handle sequence configured as a priming site, hybridization site, or binding site. In some embodiments, the barcode template is DNA, RNA, or a DNA / RNA hybrid, and the barcode sequence comprises a range of about 5 bases to about 100 bases. In some embodiments, the method for generating barcode-labeled modified targets is by amplification, hybridization, primer extension, ligation, strand transfer reaction, transposition, tagging, or combinations thereof. In some embodiments, the targets to be analyzed are selected from the group consisting of single cells, compounds, nucleic acids, proteins, microbiomes, and combinations thereof. One embodiment relates to an amplifiable single-cell sequencing method for characterizing a biological sample at the single-cell level.The method includes: providing a plurality of cells or cell nuclei from a sample, providing a plurality of barcode templates, isolating the cells or cell nuclei and more than one different barcode template in a compartment; amplifying each barcode template into a plurality of copies, and amplifying one or more classes of cell contents into a plurality of copies, wherein, in the isolated compartment, the cell contents comprise natural nucleic acid sequences or nucleic acid sequences artificially attached; coupling the amplified barcode templates with the amplified cell contents in the compartment; the amplification step and the coupling step can be carried out sequentially or simultaneously; sequencing to determine the barcode sequences in the barcode templates and their associated cell content sequences; classifying the cell contents with the same barcode sequence into a cell unit. These methods can amplify the cell contents of a single cell such that they appear as more than one cell unit during analysis. The cell contents can be DNA, RNA, proteins, lipids, organelles (located inside the cell or in the cell nucleus or associated with the outside of the cell). The cells can be eukaryotic cells and / or prokaryotic cells. The compartment can be a well, a micro-well, a droplet, a micro-droplet, a pore, and other materials capable of being isolated into different reaction units or spaces. In some embodiments, the barcode templates are oligonucleotides freely present in solution. In some embodiments, the barcode templates are encapsulated in droplets. In some embodiments, the barcoded templates are arranged in the form of nanospheres. In some embodiments, the barcode templates are immobilized on a carrier in a clonal manner (i.e., there is only one unique barcode sequence, with one or more copies) or a non-clonal manner (i.e., there are more than one unique sequences present in a single copy or multiple copies). The carrier can be a solid bead or particle, or a soluble bead or particle, or a combination thereof.
[0137] One embodiment relates to a method for sequencing a single-cell full-length transcriptome, which includes providing a plurality of cells from a biological sample; contacting the cells with a reverse transcriptase and an oligo-dT primer to generate first-strand cDNA in situ; providing a plurality of transposomes, each transposome comprising at least one transposon and a transposase; randomly in-situ tagging RNA / cDNA hybrid transcripts across the transcript; providing a plurality of barcode templates and providing amplification reagents; compartmentalizing the cells, barcode templates, and amplification reagents to generate two or more compartments, where each compartment contains a cell, one or more barcode templates having different barcode sequences, and amplification reagents; amplifying the barcode templates and tagged RNA / cDNA fragments, attaching the barcode sequences to the cDNA fragments or fragments generated from the cDNA such that a plurality of barcode-attached fragments share the same one or more barcode sequences present in the compartment; collecting the barcode-attached fragments; sequencing the barcodes and the barcoded nucleic acids to characterize the full-length transcriptome profile on a single-cell basis. In some embodiments, nuclear samples are used instead of cell samples for the method. In some embodiments, as part of the procedure, the biological sample is treated with a fixative and / or a permeabilizing agent.
[0138] Although the present disclosure has been explained with respect to one or more embodiments, it should be understood that many other possible modifications and variations can be made without departing from the spirit and scope of the present disclosure as described herein.
[0139] Furthermore, generally speaking with respect to the processes, systems, methods, etc. described herein, it should be understood that although the steps of such processes, etc. are described as occurring in a certain order, such processes can implement the steps in an order other than the order described herein. It should also be understood that certain steps can be performed simultaneously, other steps can be added, or certain steps described herein can be omitted. In other words, the description of the process herein is provided to illustrate certain embodiments and should not be construed as limiting the subject matter of the claims.
[0140] In addition, it should be understood that the foregoing description is illustrative and not restrictive. Many embodiments and applications will be apparent to those skilled in the art after reading the foregoing description in addition to the provided embodiments. When determining the scope of the present disclosure and the foregoing embodiments, reference should not be made to the foregoing description, but to the appended claims and the full scope of the equivalents given by those claims. Future developments will occur in the technologies discussed herein, and it is expected that the disclosed systems and methods will be incorporated into such future embodiments. In summary, it should be understood that the present disclosure and the described embodiments are capable of modification and variation and are limited only by the following claims.
[0141] Finally, all defined terms used in this application are intended to be given the broadest reasonable interpretation consistent with the definitions provided herein. All undefined terms used in the claims will be given the broadest reasonable interpretation in accordance with their ordinary meaning as understood by one of ordinary skill in the art, unless there is a clear contrary indication herein. Specifically, singular articles such as "a", "an", "the", etc. shall be understood to recite one or more of the recited elements, unless the claim clearly indicates a contrary limitation.
[0142] Unless otherwise indicated, the practice of the present invention employs conventional techniques of molecular biology (including recombinant techniques), microbiology, cell biology, biochemistry, and immunology within the skill of one of ordinary skill in the art. These techniques are well explained in the literature, such as "Molecular Cloning: A Laboratory Manual", second edition (Sambrook, 1989); "Oligonucleotide Synthesis" (Gait, 1984); "Animal Cell Culture" (Freshney, 1987); "Methods in Enzymology", "Handbook of Experimental Immunology" (Weir, 1996); "Gene Transfer Vectors for Mammalian Cells" (Miller and Calos, 1987); "Current Protocols in Molecular Biology" (Ausubel, 1987); "PCR: The Polymerase Chain Reaction", (Mullis, 1994); "Current Protocols in Immunology" (Coligan, 1991). These techniques can be applied to the production of the polynucleotides described herein and, accordingly, can be considered in formulating and implementing the disclosures and embodiments described herein. Particularly useful techniques for specific embodiments will be discussed in the following sections.
[0143] The following examples are put forward to provide a complete disclosure and description to those of ordinary skill in the art on how to make and use the methods of the present disclosure, and are not intended to limit the scope of the embodiments described herein. Examples
[0144] Example 1: A Scalable Method for Single-Cell Barcoding
[0145] This example describes a scalable method for barcoding the 3' end of the transcriptome at single-cell resolution, which can process thousands of cells simultaneously ( Figure 11 ).
[0146] Human HEK293 cells and mouse NIH-3T3 cells (ATCC, Manassas, VA) were cultured in Dulbecco's Modified Eagle Medium (DMEM) medium (Thermo Fisher Scientific, Waltham, MA) containing 10% fetal bovine serum (FBS) (Thermo Fisher Scientific, Waltham, MA), supplemented with 1:100 MEM non-essential amino acids (Thermo Fisher Scientific, Waltham, MA) and 1:100 penicillin / streptomycin (Thermo Fisher Scientific, Waltham, MA). After reaching 50 - 80% confluence, the cells were treated with trypsin-EDTA solution (Thermo Fisher Scientific, Waltham, MA) for 1 - 2 minutes to harvest the cells. After dilution with medium containing FBS, the cells were washed once with 1x phosphate-buffered saline (PBS) and then counted using a Countess-3 automated cell counting system (Thermo Fisher Scientific, Waltham, MA). In this experiment, approximately 250,000 HEK293 cells and 250,000 mouse NIH-3T3 cells were mixed (ratio 1:1) and processed in a low-binding 1.5 mL tube. After centrifugation (300 x g, 3 minutes), the cells were processed to stabilize RNA. Specifically, the human and mouse cell mixture was gently fixed with a mild fixative in 100 μL of 1x PBS at room temperature for 45 minutes. Subsequently, the cells were gently permeabilized with a 100 μL PBS solution containing a non-ionic detergent mixture at room temperature for 10 minutes. All reactions were carried out in the presence of RNase and protease inhibitors, and the centrifugation steps were carried out in a refrigerated centrifuge at 400 x g for 2 minutes. After the cell washing step in 100 μL solution, a filtration step was performed using a Flowmi 40 μm cell filter (Sigma-Aldrich) to remove aggregates. After the cell counting step, 50,000 fixed and permeabilized cells were incubated with reverse transcriptase (RT), primer poly-dT oligonucleotides, and dNTPs in RT buffer in a thermocycler for 30 minutes to synthesize cDNA (RT program: 10 minutes at 50°C, 3 cycles, each cycle with 12 seconds at 8°C, 45 seconds at 15°C, 45 seconds at 20°C, 30 seconds at 30°C, 2 minutes at 42°C, 2.5 minutes at 50°C, and the last step 5 minutes at 50°C). After one cell washing step in 100 μL, the intracellular cDNA molecules were tagged with a transposome in 20 μL at 37°C for 20 minutes. After another two cleanup steps and another cell counting step, approximately 12,500 cells were mixed with barcode templates and PCR reagents in a volume not exceeding 40 μL (adjusted with wash buffer).Mix the cell solution with 160 μL of the emulsion (0.2 mL barcoding reaction). In an independent experiment (different batches of human:mouse cells), mix approximately 10,000 cells with the barcode template and PCR reagents in a total volume of 300 μL, and mix the aqueous solution with 700 μL of the emulsion (1.0 mL barcoding reaction). Under controlled pipetting conditions (50 pipetting iterations), aspirate and dispense the two water-oil mixtures for approximately fifteen minutes to encapsulate the cells and barcoding reagents into droplets. The targeted ratio of the number of barcode templates to the expected number of droplets is 3 to 1 such that approximately 95% of the droplets contain at least one barcode template. The emulsion with encapsulated cells and barcoding reagents is placed into droplets and then incubated in a thermal cycler for 2 hours for barcode template amplification and cDNA barcoding (PCR program: 5 minutes at 72 °C, 30 seconds at 98 °C, 20 cycles, each cycle: 20 seconds at 98 °C, 30 seconds at 59 °C, 20 seconds at 72 °C, 5 cycles, each cycle 20 seconds at 98 °C, 2 minutes at 40 °C, 30 seconds at 72 °C, and a final step of 3 minutes at 72 °C). Then incubate the treated emulsion with 90 μL (0.2 mL reaction) or 450 μL (1.0 mL reaction) of the breaking (demulsifying) solution and vortex for 5 seconds. Separate the oil and cell debris (top layer) from the soluble molecules by centrifuging at 10,000 rpm for 5 minutes. Slowly transfer 125 μL or 625 μL of the aqueous phase to new tubes respectively. After purifying the beads with 130 μL of MagBio magnetic beads (MagBio Genomics), elute the barcoded cDNA fragments in 40 μL of low TE buffer. In addition to the PCR reagents, indexing and sequencing primers are added to the solution to generate an Illumina-compatible library (PCR program: 30 seconds at 98 °C, 8 cycles, each cycle 20 seconds at 98 °C, 30 seconds at 62 °C, and 40 seconds at 72 °C, and a final cycle of 2 minutes at 72 °C). After a new purification step using MagBio magnetic beads (0.9x), quantify and size the final library using a 4200 TapeStation system and high-sensitivity D1000 reagents (Agilent, La Jolla, California). The average size and concentration of the library are 414 base pairs (bp) and 10 mM respectively.
[0147] The library was sequenced as single-end using the NextSeq system (Illumina, San Diego, CA). Sequencing configuration: Read 1, single-end reads for 90 cycles (transcript); Index 1 (i7), 8 cycles (sample index); Index 2 (i5), 20 cycles (barcode template). Sequencing depth: The total number of reads for the 0.2 mL reaction was 103,412,571 (91.2% of the reads mapped to the genome), and the total number of reads for the 1.0 mL reaction was 103,298,991 (84.7% of the reads mapped to the genome).
[0148] After converting the sequencing data from bcl to fastq and performing demultiplexing, error correction was performed on the barcode templates, adapter sequences were trimmed, and duplicate reads were removed. For barcode template grouping, multiple barcode templates captured from the same cell contents were estimated and integrated. The resulting reads were mapped to a mixture of the reference human and mouse genomes (hg38 and Mm10) using Cell Ranger v5.0.1 software (10x Genomics), and cells were distinguished from the background using a ranked plot based on the same software: In the 0.2 mL experiment, 10,099 cells were estimated (5,149 human cells and 5,298 mouse cells; the fraction of reads in cells was 80.3%; the average number of reads per cell was 10,240; the median number of human genes per cell was 2,337; the median number of mouse genes per cell was 2,028; the total number of human genes was 28,203; the total number of mouse genes was 20,339), and in the 1.0 mL experiment, 6,715 cells were estimated (4,035 human cells and 2,699 mouse cells; the fraction of reads in cells was 68.1%; the average number of reads per cell was 15,383; the median number of human genes per cell was 1,181; the median number of mouse genes per cell was 1,933; the total number of human genes was 27,789; the total number of mouse genes was 19,755).
[0149] To verify the single-cell behavior in the two experiments ( Figure 11) The expression output generated by Cell Ranger was processed using Seurat v4.0 software (developed by the Satija lab at the New York Genome Center, New York University). Based on the human and mouse read counts, cells were visualized and most cells were distributed along the axis, which is consistent with the unique single-cell characteristics of humans or mice. Only a relatively small proportion of cells (estimated to be 3.91% (1101) in the 0.2 mL reaction and 0.48% (1102) in the 1.0 mL reaction) could be attributed to the collision between human and mouse cells in the same droplet. Therefore, the collision rate of the 0.2 mL experiment (10,099 cells) was estimated to be 7.48%, and that of the 1.0 mL experiment (6,715 cells) was 0.95%. This difference in collision rates indicates that cell collision depends on the barcoding reaction volume. The scalability of this reaction can be used to reduce the collision rate or increase the throughput of the 1 mL barcoding reaction to up to 62,500 cells.
[0150] To further verify the single-cell behavior of the barcoding reaction, in the 0.2 mL experiment, the expression of representative human genes (1103 and 1104) or mouse genes (1105 and 1106) was highlighted on the t-SNE plot. ( Figure 11 ) Overall, the expression patterns of these two genes were mutually exclusive across the cell population. Importantly, whether the multiple barcode templates were ungrouped (1103 and 1105) or cells were inferred by computationally grouping the barcode templates first (1104 and 1106), these patterns were hardly distinguishable. This observation indicates that the barcode grouping process reconstructs the cell content without inferring a large number of artificial human:mouse cells. Additionally, according to the UMAP plot, the estimated co-encapsulated human:mouse cell ratios were relatively similar whether the barcoding grouping process was skipped or this step was performed, at 4.55% and 3.38% respectively (1103 vs 1104 or 1105 vs 1106).
[0151] Example 2: 3'single-cell RNA-seq analysis of a sample containing multiple human cells (PBMC) extracted from peripheral blood
[0152] This example describes a method for barcoding the 3' end of the transcriptome at single-cell resolution, which can identify multiple cell types in a human PBMC sample from peripheral blood ( Figure 12 ).
[0153] Approximately 10 million cryopreserved PBMCs (AllCells, Alameda, California) were gently thawed, and 1 million cells were processed following the method in Example 1 after the cell harvesting step. Then libraries were also generated and sequenced following the method described in Example 1 (reaction volume 0.2 mL). Sequencing depth: 120,326,303 reads.
[0154] As described in Example 1, barcoded sorting plots were used to distinguish cells from the background: After grouping the barcode templates, an estimated 8,870 cells were identified (cell read fraction, 87.6%; average reads per cell, 8,063; median number of genes per cell, 820), or, when skipping the barcode template grouping process, an estimated 20,723 cell-associated barcodes were identified (cell read fraction, 85.3%; average reads per cell, 4,827; median number of genes per cell, 612).
[0155] Figure 12 UMAP visualization of 3' single-cell RNA-seq data is shown. The figure illustrates two analysis methods: One method is to estimate cells based on grouping multiple barcodes with similar compartment contents (1201); the other method is based on individual barcode templates without going through this barcode grouping process (1202). The identification of expected PBMC types after barcode grouping is highlighted in the figure (1203: B cells, plasma B cells, classical monocytes, non-classical monocytes, T cells, NK cells, and rare cell populations such as plasmacytoid dendritic cells or pDC cells, 0.18%, erythroid cells, 0.1%). Notably, the analysis using ungrouped barcodes (1204) has higher resolution than the analysis using grouped barcodes, and more rare cell populations can be identified using the same data (including proliferating T cells, macrophages, stimulated monocytes, and platelets; the latter accounting for no more than 0.04% of the total number of detected cell-associated barcodes). Further supporting the higher resolution of the analysis method based on ungrouped barcodes is that ungrouped barcodes can more clearly distinguish two major monocyte populations (classical cells and non-classical cells), which can be better observed when highlighting the expression of two cell type-specific gene markers (VCAN and TCFL2) in the UMAP plots generated using both analysis methods. VCAN is a marker for non-classical monocytes (1205 for barcode merged and 1207 for barcode unmerged), while TCFL2 is a marker for classical monocytes (1206 for barcode merged and 1208 for barcode unmerged).
[0156] Example 3: Full-length single-cell RNA-seq analysis of human Jurkat cells
[0157] This example describes a method for barcoding full-length transcripts at single-cell resolution ( Figure 13 ).
[0158] Human Jurkat cells (ATCC, Manassas, VA) were cultured in DMEM medium (Thermo Fisher Scientific, Waltham, MA) containing 10% FBS (Thermo Fisher Scientific, Waltham, MA), supplemented with 1:100 MEM non-essential amino acids (Thermo Fisher Scientific, Waltham, MA) and 1:100 penicillin / streptomycin (Thermo Fisher Scientific, Waltham, MA). When the cells reached a confluence of 500,000 cells per mL, the cells were collected by centrifugation and washed with 1xPBS. Approximately 500,000 Jurkat cells were processed according to Example 1. Then, a library was generated and sequenced (reaction volume: 0.2 mL) also according to the method described in Example 1. The main difference was that random hexamers were added as primer oligomers in the RT reaction, and the transposase activity of two different assembled sequences (not just one Tn5A), namely Tn5A and Tn5B, was used in the cDNA tagging step. Sequencing depth: 46,281,274 reads.
[0159] As described in Example 1, barcode sorting plots were used to distinguish cells from background: After grouping the barcode templates, the estimated number of cells was 1,526 (reads in cells, 73.8%; average reads per cell, 30,328; median number of genes per cell, 1,624). The sequencing reads were also processed as aggregates from all cells (so-called 'pseudo-bulk' analysis) and visualized as read density tracks using the University of California, Santa Cruz (UCSC) Genome Browser.
[0160] Figure 13 Shown are the pseudo-bulk read densities traced along a representative gene using the 3'-end (1301) and full-length (1302) cDNA priming methods, as well as single / double tagging (Tn5A / Tn5A&Tn5B), respectively. The figure shows that when the library was processed using the 3'scRNA-seq method (1301), the read coverage mainly concentrated at the 3'-end of the annotated genes; in contrast, when the library was processed using the full-length scRNA-seq method (1302), the read coverage spanned most of the annotated exons. The selected gene had at least three annotated isoforms (1303): isoforms 1-3 (1304). Although all three isoforms (1305) shared most of the exons and were densely covered by reads, a few isoform-specific exons were hardly covered by any reads, indicating that the isoforms containing these exons (isoforms 2 and 3 (1306)) were expressed at low levels or not expressed.
[0161] Example 4: Microbial single-cell genome analysis of a simulated mixture of five different bacterial species
[0162] This example describes a method for barcoding DNA fragments under random genomic regions at single-cell resolution for taxonomy Figure 14 ).
[0163] Five reference bacterial strains (purchased from ATCC) were cultured separately in saturated LB broth and mixed at a ratio of 1:1:1:1:1, and then cell permeabilization was performed (simulating 5, three Gram-negative cells and two Gram-positive cells): Escherichia coli (-), Bacillus subtilis (+), Citrobacter freundii (-), Klebsiella aerogene (-), and Staphylococcus epidermis (+). After mixing, the cells were washed with 1x PBS (centrifuged at 600 x g for 5 minutes at room temperature, swing bucket rotor). 10 M cells (quantifying absorbance at OD600) were gently fixed at room temperature for 45 minutes and washed twice with 1x PBS before permeabilization. On ice, the cells were permeabilized with 0.04% Tween-20 for 3 minutes. After centrifuging the cells (600 x g, 5 minutes), they were further permeabilized with 4 μg lysostaphin (Sigma-Aldrich) and 10 μg lysozyme (Sigma-Aldrich) at 37 °C for 30 minutes. Before centrifuging again, cold 1x PBS was added and the cells were pelleted by centrifuging at 600 x g for 5 minutes. After two additional washing steps in cold PBS, a tagging reaction was performed with the Tn5A and Tn5B transposome mixture at 37 °C for 1 hour. Cell encapsulation was performed using 30,000 - 250,000 cells and processed as described in Example 1.
[0164] Without the aid of a reference genome, using multiple read sequences and their barcodes, Figure 14Unsupervised hierarchical clustering of the display read sequences separates multiple barcodes by bacterial origin (1401). Briefly, using the annotations of five bacterial genomes, barcodes are distinguished according to the origin of the associated reads (genomic content). Specifically (1401), barcodes mainly containing reads of Klebsiella aerogene are clustered together; barcodes mainly containing reads of Staphylococcus epidermis are clustered together; barcodes mainly containing reads of Bacillus subtilis are clustered together; barcodes mainly containing reads of Escherichia coli are clustered together; and barcodes mainly containing reads of Citrobacter freundii are clustered together. As shown, content-based barcode clustering isolates barcodes by species, which supports that the barcoding method described herein can capture taxonomic information at single-cell resolution. In addition, different cell permeabilization and encapsulation conditions consistently indicate that, except for Bacillus subtilis and sometimes Staphylococcus epidermis, the abundance of each species can be estimated to some extent, which is consistent with the expectation based on its Gram-positive bacterial identity (1402). In summary, these results suggest that the barcoding method described herein can be used to process bacterial cells for bacterial taxonomic identification, and if efficiently permeabilized, these cells can also be quantitatively analyzed. References
[0165] Adey A. et al 2010. Genome Biol. 11, R119.
[0166] Amini S. et al 2014. Nature Genetics, 46(12):1343-1349.
[0167] Au, T. et al 2004. EMBO J., 23:3408-3420.
[0168] Buenrostro J.D. et al 2013. Nature Methods, 10(12):1213–1218.
[0169] Buenrostro, J.D. et al 2015. Nature, 523:486–490.
[0170] Burton B.M. and Baker T.A. 2003. Chemistry & Biology 10:463-472.
[0171] Caruccio N. 2011. Methods Mol. Biol. 733:241–255.
[0172] Kavanagh I, Kiiskinen L.L. and Haakana H. 2013. US Patent Application Publication US2013 / 0023423.
[0173] Kurihara K. et al. 2011. Nat. Chem. 3:775–781.
[0174] Laouini A. et al. 2012. Colloid Sci. Biotechnol. 1:147–168.
[0175] Mizuuchi M., Baker T.A. and Mizuuchi K. 1992. Cell 70, 303–311.
[0176] Savilahti H., P.A. Rice, and K. MiZuuchi. 1995. EMBO J. 14:4893-4903.
[0177] Stoeckius M., et al. 2017. Nature Methods 14:865–868.
[0178] Surette M., Buch S.J. and Chaconas G. 1987. Cell 70:303-311.
[0179] Reznikoff W.S. 2008. Annual Review of Genetics 42(1):269-286.
Claims
1. A single-cell sequencing method for characterizing a biological sample at the single-cell level, the method comprising: a) isolating a plurality of cells or a plurality of cell nuclei into compartments, wherein each cell or cell nucleus is isolated into a separate compartment having a plurality of barcode templates, each barcode template containing a barcode sequence, and wherein at least some of the compartments contain more than one group of barcode templates, each group of barcode templates having a unique barcode sequence different from other groups of barcode templates; b) amplifying at least one type of cell content in each cell or cell nucleus into multiple copies and fragmenting the cell content in each compartment into multiple fragments; c) attaching the barcode templates to the respective fragments; d) collecting the fragments to which the barcode templates are attached; and e) sequencing the fragments to which the barcode templates are attached and classifying the fragments having the same barcode sequence as belonging to the same cell unit.
2. A single-cell sequencing method for characterizing a biological sample at the single-cell level, the method comprising: a) isolating a plurality of cells or a plurality of cell nuclei and a plurality of barcode templates into compartments, wherein each cell or cell nucleus is isolated into a separate compartment having at least one barcode template containing a barcode sequence, and wherein at least some of the compartments contain at least two different barcode templates, each different barcode template having a different barcode sequence; b) amplifying at least one type of cell content in each cell or cell nucleus into multiple copies, fragmenting the cell content in each compartment into fragments, and amplifying the at least one barcode template in each compartment; c) attaching the barcode templates to the respective fragments; d) collecting the fragments to which the barcode templates are attached; and e) sequencing the fragments to which the barcode templates are attached and classifying the fragments having the same barcode sequence as belonging to the same cell unit.
3. The method according to claim 1 or 2, wherein each barcode template is a nucleotide sequence capable of functioning as a unique identifier.
4. The method according to claim 1 or 2, wherein each barcode template freely exists in solution.
5. The method according to claim 1 or 2, wherein each barcode template is immobilized on a carrier.
6. The method according to claim 5, wherein the carrier is a solid bead or particle, a soluble bead or particle, or a combination thereof.
7. The method according to claim 1 or 2, wherein the type of cell content is RNA, DNA, RNA / DNA hybrid, protein, metabolite, ligand, compound, drug, macromolecule, or a combination thereof.
8. The method according to claim 1 or 2, wherein the type of cell content is RNA, DNA, RNA / DNA hybrid, or a combination thereof.
9. The method according to claim 1 or 2, wherein the fragments are directly attached to the barcode templates.
10. The method according to claim 1 or 2, wherein the fragments are indirectly attached to the barcode templates.
11. The method according to claim 10, wherein the fragment is attached to an adaptor oligomer or linker, and the adaptor oligomer or linker is attached to the barcode template.
12. The method according to claim 1 or 2, wherein the cellular content is endogenous.
13. The method according to claim 1 or 2, wherein the cellular content is exogenous.
14. The method according to claim 1 or 2, wherein the compartment comprises a cell or nucleus without further compartmentalization; a tube or microtube; a pore or micropore; a plate; a well in a multi-well plate; a slide; a spot on a slide; a droplet; a conduit; a channel; a bottle; a chamber; or a flow cell.
15. The method according to claim 1 or 2, wherein steps (b) and (c) occur substantially simultaneously.
16. The method according to claim 1 or 2, further comprising: identifying barcode sequences attached to cellular contents derived from the same cell or nucleus, and pooling cell units corresponding to barcode sequences identified as attached to cellular contents derived from the same cell or nucleus.
17. The method according to claim 1 or 2, wherein the cell is a eukaryotic cell, a prokaryotic cell, or a combination thereof.
18. A method for single-cell transcriptome sequencing, the method comprising: a) generating cDNA from cellular RNA or nuclear RNA of cells or nuclei in a plurality of cells or nuclei; b) randomly tagging the generated cDNA along the full length of the cDNA in each cell or nucleus using a plurality of transposomes to form a plurality of tagged cDNA fragments, wherein each transposome comprises at least one transposon and a transposase; c) isolating the plurality of cells or nuclei into compartments, wherein each cell or nucleus is isolated into a separate compartment having a plurality of barcode templates, and each barcode template comprises a barcode sequence; d) attaching the barcode template to each tagged cDNA fragment in the compartment; e) collecting the barcode-attached cDNA fragments; f) sequencing the barcode and the barcode-attached cDNA fragments to characterize the transcriptome profile of each cell or nucleus on a single-cell basis.
19. A method for single-cell transcriptome sequencing, the method comprising: a) generating cDNA from cellular RNA or nuclear RNA of cells or nuclei in a plurality of cells or nuclei; b) randomly tagging the generated cDNA along the full length of the cDNA in each cell or nucleus using a plurality of transposomes to form a plurality of tagged cDNA fragments, wherein each transposome comprises at least one transposon and a transposase; c) isolating the cells or nuclei from a plurality of barcode templates, wherein each cell or nucleus is isolated into a separate compartment having at least one barcode template; d) attaching the barcode template to each tagged cDNA fragment; e) collecting the barcode-attached cDNA fragments; f) sequencing the barcode and the barcode-attached cDNA fragments to characterize the transcriptome profile of each cell on a single-cell basis.
20. The method according to claim 18, wherein the plurality of barcode templates in each compartment comprises at least two barcode template groups, and each barcode template group has a different barcode sequence.
21. The method according to claim 20, wherein the attachment generates at least two cDNA fragment groups, and each cDNA fragment group is attached to a different barcode template group.
22. The method according to claim 19, wherein the at least one barcode template is at least two different barcode templates, and each barcode template has a different barcode sequence.
23. The method according to claim 18 or 19, wherein the generated cDNA is first-strand cDNA and forms a DNA / RNA hybrid with cellular RNA or nuclear RNA.
24. The method according to claim 18 or 19, wherein the generated cDNA is first-strand and second-strand cDNA and forms double-stranded DNA.
25. The method according to claim 18 or 19, wherein the generated cDNA contains a transcript that contains the 3' end and the 5' end of cellular RNA or nuclear RNA.
26. The method according to claim 18 or 19, wherein the transcriptome profile includes the 3' end and the 5' end of cellular RNA or nuclear RNA.
27. The method according to claim 22, wherein the sequence of the cDNA fragment attached with the barcode template is converted into a full-length RNA sequence.
28. The method according to claim 18 or 19, wherein attaching the barcode template to the tagged cDNA fragment comprises amplifying the barcode template and / or amplifying the tagged cDNA fragment.
29. The method according to claim 28, wherein the amplification of the barcode template and the amplification of the tagged cDNA fragment are performed separately.
30. The method according to claim 28, wherein the amplification of the barcode template and the amplification of the tagged cDNA fragment are performed simultaneously.
31. The method according to claim 19, wherein at least one barcode template in each compartment is a single barcode template.
32. The method according to claim 18, wherein the plurality of barcode templates in each compartment are multiple copies of the same barcode template.
33. The method according to claim 18 or 19, wherein each barcode template exists freely in the solution.
34. The method according to claim 18 or 19, wherein each barcode template is immobilized on a carrier.
35. The method according to claim 34, wherein the carrier is a solid bead or particle, a soluble bead or particle, or a combination thereof.
36. The method according to claim 18 or 19, wherein the compartment includes a cell or nucleus without further compartmentalization; a tube or microtube; a well or micropore; a plate; a well in a multi-well plate; a slide; a spot on a slide; a droplet; a tube; a channel; a bottle; a chamber; or a flow cell.
37. The method according to any one of claims 1-36, wherein the cell or nucleus, or the plurality of cells or nuclei, is obtained from a biological sample or a cell culture.