High-throughput multi-omics sample analysis

Molecular barcoding of DNA fragments using a transposome enables high-throughput multi-omics analysis of single cells, addressing the need for efficient decoding of gene expression profiles and structural DNA information.

JP7814363B2Active Publication Date: 2026-02-16BECTON DICKINSON & CO
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023213088
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-05-03
Filing Date
2023-12-18
Publication Date
2026-02-16
Estimated Expiration
2039-05-01

AI Technical Summary

Technical Problem

Current methods for multi-omics analysis of single cells, such as single-cell transcriptomics and proteomics, lack efficient techniques for decoding gene expression profiles and identifying specific DNA sequences associated with cellular structures.

Method used

A method involving molecular barcoding of double-stranded DNA fragments using a transposome to generate barcoded DNA fragments, followed by sequencing and denaturing to obtain single-stranded DNA fragments, which allows for determining information about the DNA sequence and structure, including chromatin accessibility and methylation states.

Benefits of technology

Enables high-throughput analysis of multi-omics data from single cells, providing detailed insights into genomic information, chromatin accessibility, and methylome information with enhanced accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007814363000001
    Figure 0007814363000001
  • Figure 0007814363000002
    Figure 0007814363000002
  • Figure 0007814363000003
    Figure 0007814363000003
Patent Text Reader

Abstract

To provide methods and kits for sample analysis, in particular single cell analysis.SOLUTION: A method of sample analysis comprises: contacting double-stranded deoxyribonucleic acid (dsDNA) from a cell with a transposome to generate a plurality of overhang dsDNA fragments each comprising two copies of the 5' overhangs; generating a plurality of barcoded DNA fragments using a plurality of barcodes; generating a plurality of barcoded targets using a plurality of barcodes; obtaining sequencing data of the barcoded targets; detecting sequences of the plurality of barcoded DNA fragments; and determining information relating the dsDNA sequences to the structure comprising dsDNA, based on sequences of the plurality of barcoded DNA fragments in the sequencing data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 666,483, filed May 3, 2018, which is incorporated herein by reference in its entirety.

[0002] The present disclosure relates generally to the field of molecular biology, and more particularly to multi-omics analysis of cells using molecular barcoding. [Background technology]

[0003] Methods and techniques such as molecular barcoding are useful for single-cell transcriptomics analysis, which involves decoding gene expression profiles to determine the state of a cell, for example, using reverse transcription, polymerase chain reaction (PCR) amplification, and next-generation sequencing (NGS). Molecular barcoding is also useful for single-cell proteomics analysis. Methods and techniques for multi-omics analysis of single cells are needed. Summary of the Invention

[0004] Disclosed herein are embodiments of methods for sample analysis. For example, sample analysis can include, consist essentially of, or consist of single-cell analysis. In some embodiments, the method includes contacting double-stranded deoxyribonucleic acid (dsDNA) (e.g., genomic DNA (gDNA) from a cell, whether the gDNA is in the cell, in an organelle such as a nucleus or mitochondria, or in a cell fraction or extract during contact) with a transposome. The transposome can include a double-stranded nuclease (e.g., a transposase) configured to induce double-stranded DNA breaks in a structure containing the dsDNA, and two copies of an adaptor with a 5' overhang that includes a capture sequence to generate multiple overhanging dsDNA fragments, each containing two copies of the 5' overhang. The method can include barcoding a plurality of overhanging dsDNA fragments with a plurality of barcodes to generate a plurality of barcoded DNA fragments, wherein each of the plurality of barcodes comprises a cell label sequence, a molecular label sequence, and a capture sequence, and wherein at least two of the plurality of barcodes comprise different molecular label sequences and at least two of the plurality of barcodes comprise the same cell label sequence. The method can include detecting the sequences of the plurality of barcoded DNA fragments. The method can also include determining information relating the dsDNA sequence to a structure comprising the dsDNA based on the sequences of the plurality of barcoded DNA fragments in the sequencing data. The method can further include contacting the plurality of overhanging dsDNA fragments with a polymerase to generate a plurality of complementary dsDNA fragments, each comprising a sequence complementary to at least a portion of the 5' overhang; and denaturing the plurality of complementary dsDNA fragments to generate a plurality of single-stranded DNA (ssDNA) fragments, wherein the ssDNA fragments are barcoded, thereby barcoding the DNA fragments. In some embodiments, the dsDNA comprises, consists essentially of, or consists of gDNA.In any of the methods of sample analysis described herein, the transposome may target a specific structure containing dsDNA, such as chromatin, a specific DNA methylation state, DNA in a specific organelle, etc. It is contemplated that the method of sample analysis may identify a specific DNA sequence associated with the structure targeted by the transposome, such as chromatin-accessible DNA, construct DNA, organelle DNA, etc.

[0005] In some embodiments, a method of analyzing a sample includes generating a plurality of nucleic acid fragments from dsDNA (e.g., gDNA from a cell, regardless of whether the gDNA is in contact with, within, or in the nucleus of the cell), each of the plurality of nucleic acid fragments comprising a capture sequence, a complement of the capture sequence, a reverse complement of the capture sequence, or a combination thereof; barcoding the plurality of nucleic acid fragments using a plurality of barcodes to generate a plurality of barcoded DNA fragments, each of the plurality of barcodes comprising a cell label sequence, a molecular label sequence, and a capture sequence, wherein at least two of the plurality of barcodes comprise different molecular label sequences and at least two of the plurality of barcodes comprise the same cell label sequence; and detecting the sequence of the plurality of barcoded DNA fragments. The method may further include determining information relating the dsDNA sequence to a structure comprising the dsDNA based on the sequences of the plurality of barcoded DNA fragments in the sequencing data.

[0006] In some embodiments, for any of the methods of sample analysis described herein, generating a plurality of nucleic acid fragments may include contacting the dsDNA with a transposome, the transposome including a double-stranded nuclease configured to induce double-stranded DNA breaks in a structure containing the dsDNA, and two copies of an adapter including a capture sequence, to generate a plurality of complementary dsDNA fragments, each of which includes a sequence complementary to the capture sequence. The double-stranded nuclease may be loaded with two copies of the adapter. The method may further include denaturing the complementary dsDNA fragments to generate a plurality of single-stranded DNA (ssDNA) fragments. The method may include barcoding the plurality of ssDNA fragments, thereby generating a plurality of barcoded DNA fragments. The method may further include denaturing the barcoded DNA fragments to generate barcoded single-stranded DNA (ssDNA) fragments.

[0007] In some embodiments, for any of the methods of sample analysis described herein, generating a plurality of nucleic acid fragments may include contacting the dsDNA with a transposome, the transposome including a double-stranded nuclease configured to induce double-stranded DNA cleavage in a structure containing the dsDNA, and two copies of an adapter having a 5' overhang comprising a capture sequence, to generate a plurality of overhanging dsDNA fragments, each having two copies of the 5' overhang; and contacting the plurality of overhanging dsDNA fragments with a polymerase to generate a plurality of complementary dsDNA fragments, each comprising a sequence complementary to at least a portion of the 5' overhang. The double-stranded nuclease may be loaded with two copies of the adapter. The method may further include denaturing the complementary dsDNA fragments to generate a plurality of single-stranded DNA (ssDNA) fragments. The method may further include attaching barcodes to the ssDNA fragments, thereby generating barcoded DNA. The method may further include denaturing the barcoded DNA fragments to generate barcoded single-stranded DNA (ssDNA) fragments. In some embodiments, for any of the methods of sample analysis described herein, the barcoded DNA fragments may be ssDNA fragments.

[0008] In some embodiments, for any of the methods of sample analysis described herein, none of the plurality of complementary dsDNA fragments comprises an overhang (e.g., a 3' overhang or a 5' overhang). In some embodiments, for any of the methods of sample analysis described herein, the adapter may comprise a DNA end sequence of a transposon. By way of example, the double-stranded nuclease configured to induce double-stranded DNA breaks in the dsDNA-containing structure may comprise a transposase such as Tn5 transposase. Examples of other suitable transposases are described herein. In some embodiments, for any of the methods of sample analysis described herein, each of the plurality of complementary dsDNA fragments comprises a blunt end.

[0009] In some embodiments, for any of the methods of sample analysis described herein, generating a plurality of nucleic acid fragments comprises fragmenting dsDNA to generate a plurality of dsDNA fragments. Fragmenting the dsDNA may comprise contacting the dsDNA with a restriction enzyme to generate a plurality of dsDNA fragments, each having one or two blunt ends. In some embodiments, at least one of the plurality of dsDNA fragments may comprise a blunt end. In some embodiments, at least one of the plurality of dsDNA fragments may comprise a 5' overhang and / or a 3' overhang. In some embodiments, none of the plurality of dsDNA fragments comprises a blunt end.

[0010] In some embodiments, for any of the methods of sample analysis described herein, fragmenting the dsDNA can include contacting the dsDNA with a CRISPR-associated protein (e.g., Cas9 or Cas12a) to generate a plurality of dsDNA fragments. By way of example, a guide RNA complementary to a target DNA motif or sequence can be used to target the CRISPR-associated protein to generate a double-stranded DNA break at the target DNA motif or sequence.

[0011] In some embodiments, for any of the sample analysis methods described herein, generating a plurality of nucleic acid fragments comprises adding two copies of an adaptor comprising a sequence complementary to a capture sequence to at least one of the plurality of dsDNA fragments to generate a plurality of dsDNA fragments.For example, the adaptor can be added by a transposase as described herein.For example, adding two copies of an adaptor can comprise ligating two copies of an adaptor to at least one of the plurality of dsDNA fragments to generate a plurality of dsDNA fragments comprising an adaptor.

[0012] In some embodiments, for any of the methods of sample analysis described herein, the capture sequence comprises a poly(dT) region. The sequence complementary to the capture sequence may comprise a poly(dA) region.

[0013] In some embodiments, for any of the methods of sample analysis described herein, fragmenting the dsDNA can include contacting the dsDNA with a restriction enzyme to generate a plurality of dsDNA fragments, at least one of the plurality of dsDNA fragments comprising a capture sequence. The capture sequence can be complementary to the sequence of the 5' overhang. The sequence complementary to the capture sequence can comprise the sequence of the 5' overhang. In some embodiments, the capture sequence comprises a sequence that does not contain 3, 4, 5, 6, or more consecutive T's. For example, the capture sequence can comprise a sequence unique to one or both strands of the target dsDNA.

[0014] In some embodiments, for any of the methods of sample analysis described herein, the dsDNA is located within a cellular organelle, such as a nucleus. The method may include, for example, permeabilizing the nucleus with a detergent such as Triton X-100 to generate a permeabilized nucleus. The method may include fixing the cells containing the nucleus before permeabilizing the nucleus. In some embodiments, for any of the methods of sample analysis described herein, the dsDNA is located within at least one of a nucleus, a nucleolus, a mitochondria, or a chloroplast. In some embodiments, the dsDNA is selected from the group consisting of nuclear DNA (e.g., a portion of chromatin), nucleolar DNA, genomic DNA, mitochondrial DNA, chloroplast DNA, construct DNA, viral DNA, or a combination of two or more of the listed items. Examples of construct DNA may include plasmids, cloning vectors, expression vectors, hybrid vectors, minicircles, cosmids, viral vectors, BACs, YACs, and HACs. For example, viral DNA present in extragenomic DNA may be inserted into the host genome. For example, the methods of sample analysis described herein may quantify DNA or types of DNA in one or more organelles of a cell. For example, the sample analysis method described herein can quantify viral DNA or viral load of DNA in cells. For example, the sample analysis method described herein can quantify construct DNA (e.g., plasmid, cloning vector, expression vector, hybrid vector, minicircle, cosmid, viral vector, BAC, YAC and / or HAC) in cells. Therefore, it is believed that this method can provide information about transposome-accessible structures containing dsDNA.

[0015] In some embodiments, for any of the methods of sample analysis described herein, the method includes denaturing a plurality of nucleic acid fragments to generate a plurality of ssDNA fragments, wherein barcoding the plurality of nucleic acid fragments includes barcoding the plurality of ssDNA fragments using the plurality of barcodes to generate a plurality of barcoded ssDNA fragments. In some embodiments, for any of the methods of sample analysis described herein, the adapter includes a promoter sequence. Generating the plurality of nucleic acid fragments can include transcribing the plurality of dsDNA fragments using in vitro transcription to generate a plurality of ribonucleic acid (RNA) molecules, wherein barcoding the plurality of nucleic acid fragments can include barcoding the plurality of RNA molecules. The promoter sequence can include a T7 promoter sequence.

[0016] In some embodiments, for any of the methods of sample analysis described herein, determining information about the dsDNA (e.g., gDNA) includes determining chromatin accessibility of the dsDNA (e.g., gDNA) based on the sequences and / or abundance of a plurality of barcoded DNA fragments in the obtained sequencing data. Determining chromatin accessibility of the dsDNA may include aligning the sequences of the plurality of barcoded DNA fragments with a reference sequence of dsDNA (e.g., gDNA); and identifying regions of the dsDNA corresponding to ends of barcoded DNA fragments (e.g., barcoded ssDNA fragments) of the plurality of ssDNA fragments as having accessibility above a threshold. Determining the chromatin accessibility of dsDNA (e.g., gDNA) may include aligning the sequences of a plurality of barcoded DNA fragments (e.g., ssDNA fragments) with a reference sequence of dsDNA (e.g., gDNA); and determining the accessibility of regions of dsDNA (e.g., gDNA) corresponding to the ends of the barcoded DNA fragments (e.g., barcoded ssDNA fragments) of the plurality of barcoded DNA fragments (e.g., barcoded ssDNA fragments) based on the number of barcoded DNA fragments (e.g., barcoded ssDNA fragments) of the plurality of barcoded DNA fragments (e.g., barcoded ssDNA fragments) in the sequencing data.

[0017] In some embodiments, for any of the methods of sample analysis described herein, determining information about the dsDNA (e.g., gDNA) includes determining genomic information of the dsDNA based on the sequences of multiple barcoded DNA fragments (e.g., barcoded ssDNA fragments) in the obtained sequencing data. The method of sample analysis may include digesting nucleosomes associated with the dsDNA. Determining genomic information of the dsDNA may include determining at least a partial sequence of the dsDNA by aligning the sequences of the multiple barcoded DNA fragments (e.g., barcoded ssDNA fragments) with a reference sequence of the dsDNA.

[0018] In some embodiments, for any of the methods of sample analysis described herein, determining information relating dsDNA (e.g., gDNA) to a structure containing the dsDNA includes determining methylome information of the dsDNA (e.g., gDNA) based on the sequences of multiple barcoded DNA fragments in the obtained sequencing data. The method of sample analysis may include digesting nucleosomes associated with the dsDNA. The method of sample analysis may include bisulfite conversion of cytosine bases of multiple single-stranded DNA fragments of multiple overhanging DNA fragments or multiple nucleic acid fragments (e.g., obtained by denaturing the overhanging DNA fragments or multiple nucleic acid fragments) to generate multiple bisulfite-converted ssDNA fragments having uracil bases. Barcoding the multiple overhanging DNA fragments or barcoding the multiple nucleic acid fragments may include barcoding the multiple bisulfite-converted ssDNA fragments using the multiple barcodes to generate multiple barcoded ssDNA fragments. Determining the methylome information may include determining that positions of a plurality of barcoded DNA fragments (e.g., barcoded ssDNA fragments) in the sequencing data have a thymine base and that corresponding positions in a reference sequence of dsDNA have a cytosine base, thereby determining that corresponding positions in the dsDNA have a methylcytosine base.

[0019] In some embodiments, for any of the methods of sample analysis described herein, barcoding comprises probabilistically barcoding a plurality of DNA fragments (e.g., ssDNA fragments) or a plurality of nucleic acids using a plurality of barcodes to generate a plurality of probabilistically barcoded DNA fragments. Barcoding can comprise barcoding a plurality of DNA fragments (e.g., ssDNA fragments) or a plurality of nucleic acid fragments using a plurality of barcodes associated with particles to generate a plurality of barcoded ssDNA fragments, wherein the barcodes associated with the particles comprise the same cellular label sequence and at least 100 different molecular label sequences.

[0020] In some embodiments, for any of the methods of sample analysis described herein, at least one barcode of the plurality of barcodes may be immobilized on a particle. At least one barcode of the plurality of barcodes may be partially immobilized on a particle. At least one barcode of the plurality of barcodes may be encapsulated in a particle. At least one barcode of the plurality of barcodes may be partially encapsulated in a particle. The particle may be disintegrable. The particle may comprise a rupturable hydrogel particle. The particle may comprise sepharose beads, streptavidin beads, agarose beads, magnetic beads, conjugated beads, protein A conjugated beads, protein G conjugated beads, protein A / G conjugated beads, protein L conjugated beads, oligo(dT) conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, or any combination thereof. The particles may comprise a material selected from the group consisting of polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic material, ceramic, plastic, glass, methylstyrene, acrylic polymer, titanium, latex, Sepharose, cellulose, nylon, silicone, and any combination thereof. In some embodiments, for any of the methods of sample analysis described herein, at least one barcode of a plurality of barcodes may be separated from other barcodes. It is contemplated that separating may include, for example, disposing the barcode on a solid support, such as a particle described herein, and disposing the barcode in a droplet (e.g., a microdroplet), such as a hydrogel droplet, or in a well of a substrate, such as a microwell, or in a chamber of a fluidic device (e.g., a microfluidic device).

[0021] In some embodiments, for any of the methods of sample analysis described herein, the barcode of the particle may include molecular labels having at least 1,000 different molecular label sequences. The barcode of the particle may include molecular labels having at least 10,000 different molecular label sequences. The molecular labels of the barcode may include a random sequence. The particle may include at least 10,000 barcodes.

[0022] In any of the single-cell analysis methods described herein, barcoding the plurality of overhanging DNA fragments or the plurality of nucleic acid fragments may include contacting the plurality of ssDNA fragments (of the DNA fragments or nucleic acid fragments) with capture sequences of the plurality of barcodes; and transcribing the plurality of ssDNA fragments using the plurality of barcodes to generate a plurality of barcoded ssDNA fragments. The method of sample analysis may include amplifying the plurality of barcoded ssDNA fragments to generate a plurality of amplified barcoded DNA fragments before obtaining sequencing data of the plurality of barcoded ssDNA fragments. Amplifying the plurality of barcoded ssDNA fragments may include amplifying the barcoded ssDNA fragments by polymerase chain reaction (PCR).

[0023] In some embodiments, any of the methods of sample analysis described herein may include barcoding a plurality of nuclear targets with a plurality of barcodes to generate a plurality of barcoded targets; and obtaining sequencing data for the barcoded targets.

[0024] In some embodiments, for any of the methods of sample analysis described herein, the dsDNA from the cells is selected from the group consisting of nuclear DNA, nucleolar DNA, genomic DNA, mitochondrial DNA, chloroplast DNA, structural DNA, viral DNA, or a combination of two or more of the listed items. In some embodiments, for any of the methods of sample analysis described herein, the 5' overhang comprises a poly dT sequence. In some embodiments, for any of the methods of sample analysis described herein, the method comprises capturing ssDNA fragments of a plurality of barcoded sDNA fragments on a particle comprising oligonucleotides comprising a capture sequence, a cell label sequence, and a molecular label sequence, wherein the capture sequence comprises a poly dT sequence that binds to the poly A tail on the ssDNA fragment, and the captured ssDNA fragment comprises a methylated cytidine; Bisulfite The method further includes the steps of: performing a conversion reaction to convert methylated cytidines to thymidines; extending the ssDNA fragments in a 5'-3' direction to generate barcoded ssDNA fragments containing thymidines, wherein the barcoded ssDNA comprises a capture sequence, a molecular beacon sequence, and a cell beacon sequence; extending oligonucleotides in a 5'-3' direction using a reverse transcriptase or a polymerase, or a combination thereof, to generate complementary DNA strands complementary to the barcoded ssDNA containing thymidines; denaturing the barcoded ssDNA and the complementary DNA strand to generate single-stranded sequences; and amplifying the single-stranded sequences. The method may further include determining whether positions of the plurality of ssDNA fragments in the sequencing data have thymine bases and whether corresponding positions in a reference sequence of dsDNA have cytosine bases. Bisulfite After the conversion reaction, the step of determining that the position corresponding to the thymine base in the reference sequence is a cytosine base.

[0025] In some embodiments, for any of the methods of sample analysis described herein, the double-stranded nuclease of the transposome is selected from the group consisting of a transposase, a restriction endonuclease, a CRISPR-associated protein, a double-strand-specific nuclease, or a combination thereof. In some embodiments, for any of the methods of sample analysis described herein, the transposome further comprises an antibody or fragment thereof, an apatomer, or a DNA-binding domain that binds to a structure comprising dsDNA. In some embodiments, for any of the methods of sample analysis described herein, the transposome further comprises a ligase.

[0026] In some embodiments, a nucleic acid reagent is described. The nucleic acid reagent may include a capture sequence, a barcode, a primer binding site, and a double-stranded DNA binder. The capture sequence may include a poly(A) region. The primer binding site may include a universal primer binding site. The nucleic acid reagent may be plasma membrane impermeable. In some embodiments, the nucleic acid reagent is configured to specifically bind to dead cells. In some embodiments, the nucleic acid reagent does not bind to live cells.

[0027] In some embodiments, for any of the methods of sample analysis described herein, the method further comprises contacting the cells with a nucleic acid reagent. The nucleic acid reagent may be as described herein. The nucleic acid reagent may comprise a capture sequence, a barcode, a primer binding site, and a double-stranded DNA binding agent. The cells may be dead cells, and the nucleic acid binding reagent may bind to double-stranded DNA in the dead cells. The method may include washing the dead cells to remove excess nucleic acid binding reagent. The method may include lysing the dead cells. Lysis may release the nucleic acid binding reagent. The method may include barcoding the nucleic acid binding reagent. In some embodiment methods, the cells are associated with a solid support comprising an oligonucleotide comprising a cell labeling sequence, and the barcoding comprises barcoding the nucleic acid binding reagent with the cell labeling sequence. The solid support may comprise multiple oligonucleotides each comprising a cell labeling sequence and a different molecular labeling sequence. In some embodiments, the method further comprises sequencing the barcoded nucleic acid binding reagent and determining the presence of dead cells based on the presence of the barcode on the nucleic acid reagent. In some embodiments, the method further comprises associating two or more cells with different solid supports each containing a different cell marker, whereby each of the two or more cells is associated one-to-one with a different cell marker. In some embodiments, the method further comprises determining the number of dead cells in the sample based on the number of unique cell markers associated with the barcodes of the nucleic acid reagents. Determining the number of molecular marker sequences and control barcode sequences having unique sequences associated with the cell markers may include, for each cell marker in the sequencing data, determining the number of molecular marker sequences and control barcode sequences having the highest number of unique sequences associated with the cell marker. In some embodiment methods, the nucleic acid binding reagent does not enter live cells and therefore does not bind to double-stranded DNA in live cells.In some embodiments, the method further comprises contacting the dead cells with a protein-binding reagent associated with a unique identifier oligonucleotide, wherein the protein-binding reagent binds to a protein of the dead cells; and barcoding the unique identifier oligonucleotide. In some embodiments, the protein-binding reagent comprises an antibody, a tetramer, an aptamer, a protein scaffold, an invasin, or a combination thereof. In some embodiments, the protein target of the protein-binding reagent is selected from a group comprising 10 to 100 different protein targets, or the cellular component target of the cellular component binding reagent is selected from a group comprising 10 to 100 different cellular component targets. In some embodiments, the protein target of the protein-binding reagent comprises a carbohydrate, a lipid, a protein, an extracellular protein, a cell surface protein, a cell marker, a B-cell receptor, a T-cell receptor, a major histocompatibility complex, a tumor antigen, a receptor, an integrin, an intracellular protein, or any combination thereof. In some embodiments, the protein-binding reagent comprises an antibody or fragment thereof that binds to a cell surface protein. In some embodiments, the barcoding is with a barcode comprising a molecular beacon sequence.

[0028] Some embodiments include a method of sample analysis. The method can include contacting dead cells of a sample with a nucleic acid binding reagent, the nucleic acid binding reagent comprising a capture sequence, a barcode, a primer binding site, and a double-stranded DNA binding agent. The nucleic acid binding reagent can bind to double-stranded DNA in the dead cells. The method can include removing excess nucleic acid binding reagent from the dead cells. The method can include lysing the dead cells, thereby releasing the nucleic acid binding reagent from the dead cells. The method can include attaching a barcode to the nucleic acid binding reagent. In some embodiment methods, the attaching barcode comprises capturing the dead cells on a solid support, such as a bead, the solid support comprising a cell labeling sequence and a molecular labeling sequence. In some embodiments, the method further includes determining the number of unique molecular labeling sequences associated with each cell labeling sequence and determining the number of dead cells in the sample based on the number of unique cell labeling sequences associated with the molecular labeling sequence. In some embodiments, determining the number of molecular label sequences and control barcode sequences having unique sequences associated with the cell label comprises, for each cell label in the sequencing data, determining the number of molecular label sequences having the highest number of unique sequences associated with the cell label. In some embodiments, the method further comprises contacting the dead cells with a protein-binding reagent associated with a unique identifier oligonucleotide. The protein-binding reagent may bind to a protein of the dead cell. The method may further comprise barcoding the unique identifier oligonucleotide. In some embodiments, the protein-binding reagent is associated with two or more sample-indexing oligonucleotides having the same sequence. In some embodiments, the protein-binding reagent is associated with two or more sample-indexing oligonucleotides having different sample-indexing sequences. In some embodiments, the protein-binding reagent comprises an antibody, a tetramer, an aptamer, a protein scaffold, an invasin, or a combination thereof.In some embodiment methods, the protein target of the protein binding reagent is selected from a group comprising 10 to 100 different protein targets, or the cellular component target of the cellular component binding reagent is selected from a group comprising 10 to 100 different cellular component targets. In some embodiment methods, the protein target of the protein binding reagent comprises a carbohydrate, a lipid, a protein, an extracellular protein, a cell surface protein, a cell marker, a B cell receptor, a T cell receptor, a major histocompatibility complex, a tumor antigen, a receptor, an integrin, an intracellular protein, or any combination thereof. In some embodiment methods, the protein binding reagent comprises an antibody or fragment thereof that binds to a cell surface protein.

[0029] In some embodiment methods, the capture sequence and the sequence complementary to the capture sequence are a specific pair of complementary nucleic acids at least 5 nucleotides to about 25 nucleotides in length.

[0030] In some embodiments, a method for sample analysis is described. The method may include contacting double-stranded deoxyribonucleic acid (dsDNA) from a cell with a transposome comprising a double-stranded nuclease configured to induce double-stranded DNA breaks in a structure comprising the dsDNA and two copies of an adapter having a 5' overhang comprising a capture sequence to generate a plurality of overhanging dsDNA fragments, each comprising two copies of the 5' overhang. The method may also include contacting the plurality of overhanging dsDNA fragments with a polymerase to generate a plurality of complementary dsDNA fragments, each comprising a sequence complementary to at least a portion of a respective 5' overhang. The method may also include denaturing the plurality of complementary dsDNA fragments to generate a plurality of single-stranded DNA (ssDNA) fragments. The method can include barcoding a plurality of ssDNA fragments with a plurality of barcodes to generate a plurality of barcoded ssDNA fragments, wherein each of the plurality of barcodes includes a cellular label sequence, a molecular label sequence, and a capture sequence, and at least two of the plurality of barcodes include different molecular label sequences, and the plurality of barcodes include the same cellular label sequence. The method can include obtaining sequencing data for the plurality of barcoded ssDNA fragments. The method can also include quantifying the amount of dsDNA in a cell based on the amount of unique molecular label sequences associated with the same cellular label sequence.In some embodiments, the method further comprises capturing ssDNA fragments of a plurality of ssDNA fragments on a solid support comprising oligonucleotides comprising a capture sequence, a cellular label sequence, and a molecular label sequence, wherein the capture sequence comprises a poly(dT) sequence that binds to the poly(A) tails on the ssDNA fragments; extending the ssDNA fragments in the 5'-3' direction to generate barcoded ssDNA fragments, wherein the barcoded ssDNA comprises the capture sequence, the molecular label sequence, and the cellular label sequence; extending the oligonucleotides in the 5'-3' direction using reverse transcriptase or polymerase, or a combination thereof, to generate complementary DNA strands complementary to the barcoded ssDNA; denaturing the barcoded ssDNA and the complementary DNA strand to generate single-stranded sequences; and amplifying the single-stranded sequences. In some embodiments, the method further comprises bisulfite conversion of cytosine bases of the plurality of ssDNA fragments to generate a plurality of bisulfite-converted ssDNA fragments comprising uracil bases.

[0031] In any of the methods described herein, the dsDNA may comprise a construct DNA. The construct DNA may be selected from the group consisting of a plasmid, a cloning vector, an expression vector, a hybrid vector, a minicircle, a cosmid, a viral vector, a BAC, a YAC, and a HAC. In some embodiments, the number of construct DNAs is from 1 to about 1 x 10 6 The range is.

[0032] In any of the methods described herein, the dsDNA can include viral DNA. The viral DNA load in a cell is about 1 x 10 2 ~1×10 6 The range may be:

[0033] In some embodiments, a kit for sample analysis is described. The kit may include a transposome described herein and a plurality of barcodes described herein. Each transposome may include a double-stranded nuclease (e.g., a transposase described herein) configured to induce double-stranded DNA breaks in a structure containing dsDNA and two copies of an adapter with a 5' overhang containing a capture sequence. Optionally, the transposome further includes a ligase. Each barcode may include a cellular label sequence, a molecular label sequence, and a capture sequence, such as a poly-T sequence. At least two of the plurality of barcodes may include different molecular label sequences, and at least two of the plurality of barcodes may include the same cellular label sequence. For example, the barcodes may include at least 10, 50, 100, 500, 1000, 5000, 10,000, 50,000, or 100,000 different molecular labels. The barcodes may be immobilized on particles described herein. All of the barcodes on the same particle may include the same cellular label. In some embodiment kits, the barcodes are separated into wells of the substrate, and all of the barcodes separated into each well can contain the same cell-labeling sequence, with different wells containing different cell-labeling sequences. [Brief explanation of the drawings]

[0034] [Figure 1] 1 shows a non-limiting exemplary barcode. [Figure 2] 1 illustrates a non-limiting exemplary workflow for barcoding and digital counting. [Figure 3] FIG. 1 is a schematic diagram showing a non-limiting exemplary process for creating an indexed library of barcoded targets from multiple targets. [Figure 4A] FIG. 1 shows a schematic diagram of a non-limiting exemplary method for high-throughput capture of multi-omics information from single cells. [Figure 4B] FIG. 1 shows a schematic diagram of a non-limiting exemplary method for high-throughput capture of multi-omics information from single cells. [Figure 5]5A-5B schematically illustrate a non-limiting exemplary method for capturing genomic and chromatin accessibility information from single cells with enhanced signal intensity. [Figure 6] 1A-1C are schematic diagrams illustrating non-limiting exemplary nucleic acid reagents of some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0035] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In these drawings, like numerals typically identify like elements, unless the context dictates otherwise. The exemplary embodiments described in the detailed description, drawings, and claims are not intended to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that aspects of the present disclosure, as generally described herein and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a variety of different configurations, all of which are expressly contemplated herein and form a part of this disclosure.

[0036] All patents, published patent applications, other publications, and sequences from GenBank and other databases mentioned herein are incorporated by reference in their entirety for relevant art.

[0037] Barcodes, such as stochastic barcodes with molecular labels (also called molecular beacons (MIs)) with different molecular label differences, can be used to determine the abundance of nucleic acid targets, such as their relative or absolute abundance. Stochastic barcoding can be performed using the Precise™ assay (Cellular Research, Inc., Palo Alto, CA) and the Rhapsody™ assay (Becton, Dickinson and Company, Franklin Lakes, NJ). The Precise™ or Rhapsody™ assay uses a non-depleting pool of stochastic barcodes with a large number (e.g., 6561 to 65536) of unique molecular label sequences on poly(T) oligonucleotides to hybridize all poly(A)-mRNAs in a sample during the reverse transcription (RT) step. The stochastic barcodes can contain universal PCR priming sites. During RT, target gene molecules react randomly with the stochastic barcodes. Each target molecule can hybridize with a probabilistic barcode to generate a probabilistically barcoded complementary ribonucleotide acid (cDNA) molecule. After labeling, the probabilistically barcoded cDNA molecules from the microwells of the microwell plate can be pooled into a single tube for PCR amplification and sequencing. The raw sequencing data can be analyzed to obtain the number of reads, the number of probabilistic barcodes with unique molecular label sequences, and the number of mRNA molecules.

[0038] Disclosed herein are embodiments of methods of sample analysis. For example, any of the methods of sample analysis described herein can include, consist of, or consist essentially of single-cell analysis. The methods of sample analysis can be used for multi-omics analysis using molecular barcoding (such as the Precise™ assay and the Rhapsody™ assay). In some embodiments, the methods of sample analysis include contacting double-stranded deoxyribonucleic acid (dsDNA) with a transposome, the transposome comprising a double-stranded nuclease configured to induce double-stranded DNA breaks in a structure comprising the dsDNA, and two copies of an adapter with a 5' overhang comprising a capture sequence, to generate a plurality of overhanging double-stranded DNA (dsDNA) fragments, each having two copies of the 5' overhang. The double-stranded nuclease (e.g., transposase) can be loaded with two copies of the adapter. The method may include contacting a plurality of overhanging dsDNA fragments (including 5' overhangs) with a polymerase to generate a plurality of complementary dsDNA fragments, each comprising a sequence complementary to at least a portion of the 5' overhang; denaturing the plurality of complementary dsDNA fragments (each comprising a sequence complementary to at least a portion of the 5' overhang) to generate a plurality of single-stranded DNA (ssDNA) fragments; barcoding the plurality of ssDNA fragments using a plurality of barcodes to generate a plurality of barcoded ssDNA fragments, each of the plurality of barcodes comprising a cell label sequence, a molecular label sequence, and a capture sequence, wherein at least two of the plurality of barcodes comprise different molecular label sequences and at least two of the plurality of barcodes comprise the same cell label sequence; obtaining sequencing data of the plurality of barcoded ssDNA fragments; and determining information about the dsDNA (e.g., gDNA) based on the sequences of the plurality of ssDNA fragments in the obtained sequencing data.

[0039] In some embodiments, for any of the methods of sample analysis described herein, the double-stranded DNA can comprise, consist essentially of, or consist of any double-stranded DNA, such as genomic DNA (gDNA), organelle DNA (e.g., nuclear DNA, nucleolar DNA, genomic DNA, mitochondrial DNA, and chloroplast DNA), viral DNA, and / or construct DNA (e.g., plasmids, cloning vectors, expression vectors, hybrid vectors, minicircles, cosmids, viral vectors, and / or artificial chromosomes, such as BACs, YACs, and HACs).

[0040] In some embodiments, for any of the methods of sample analysis described herein, the construct DNA is selected from the group consisting of a plasmid, a cloning vector, an expression vector, a hybrid vector, a minicircle, a cosmid, a viral vector, a BAC, a YAC, and a HAC, or a combination of two or more of any of the listed items.

[0041] In some embodiments, for any of the methods of sample analysis described herein, the number of construct DNAs is between 1 and about 1 x 10 6 The range is.

[0042] In some embodiments, for any of the methods of sample analysis described herein, the viral DNA load is about 1 x 10 2 ~1×10 6 The range is.

[0043] Several suitable double-stranded DNA binding reagents can be used in the nucleic acid reagent and sample analysis methods described herein. In some embodiments, for any nucleic acid reagent and / or method of sample analysis described herein, the double-stranded DNA acid-binding reagent is selected from the group consisting of, but not limited to, anthracyclines (e.g., aclarubicin, aldoxorubicin, amrubicin, annamycin, bohemic acid, carubicin, cosmomycin B, daunorubicin, doxorubicin, epirubicin, idarubicin, menogaril, nogalamycin, pirarubicin, sabarubicin, valrubicin, zoptarelin doxorubicin, and zorubicin), amykelin, 9-aminoacridine, 7-aminoactinomycin D, amsacrine, dactinomycin, daunorubicin, doxorubicin, ellipticine, ethidium bromide, mitoxantrone, pirarubicin, pixantrone, proflavine, and psoralen, or a combination of two or more of the listed items.

[0044] In some embodiments, any of the methods of sample analysis described herein includes generating a plurality of nucleic acid fragments from cellular double-stranded deoxyribonucleic acid (dsDNA), each of the plurality of nucleic acid fragments comprising a capture sequence, a complement of the capture sequence, a reverse complement of the capture sequence, or a combination thereof; barcoding the plurality of nucleic acid fragments using a plurality of barcodes to generate a plurality of barcoded single-stranded deoxyribonucleic acid (ssDNA) fragments, each of the plurality of barcodes comprising a cell label sequence, a molecular label sequence, and a capture sequence, wherein at least two of the plurality of barcodes comprise different molecular label sequences and at least two of the plurality of barcodes comprise the same cell label sequence; obtaining sequencing data of the plurality of barcoded ssDNA fragments; and determining information about the dsDNA based on the sequences of the plurality of ssDNA fragments in the obtained sequencing data.

[0045] Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. See, for example, Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989). For the purposes of this disclosure, information on the following terms is provided below.

[0046] As used herein, the term "adapter" has its conventional and usual meaning in the art, taking this specification into consideration. It refers to a sequence for facilitating the amplification, sequencing, and / or capture of a related nucleic acid. The related nucleic acid may include a target nucleic acid. The related nucleic acid may include one or more of a spatial label, a target label, a sample label, an index label, or a barcode sequence (e.g., a molecular label). The adapter may be linear. The adapter may be a pre-adenylated adapter. The adapter may be double-stranded or single-stranded. One or more adapters may be located at the 5' or 3' end of a nucleic acid. When an adapter includes a known sequence at the 5' and 3' ends, the known sequence may be the same or different sequences. An adapter located at the 5' and / or 3' end of a polynucleotide may be capable of hybridizing to one or more oligonucleotides immobilized on a surface. In some embodiments, the adapter may include a universal sequence. A universal sequence may be a region of nucleotide sequence common to two or more nucleic acid molecules. The two or more nucleic acid molecules may also have regions of different sequences. Thus, for example, the 5' adapters can contain identical and / or universal nucleic acid sequences, and the 3' adapters can contain identical and / or universal sequences. A universal sequence that can be present in different members of a plurality of nucleic acid molecules can enable replication or amplification of multiple different sequences using a single universal primer complementary to the universal sequence. Similarly, at least one, two (e.g., pairs), or more universal sequences that can be present in different members of a collection of nucleic acid molecules can enable replication or amplification of multiple different sequences using at least one, two (e.g., pairs), or more single universal primers complementary to the universal sequence. Thus, a universal primer comprises a sequence that can hybridize to such a universal sequence. A target nucleic acid sequence-bearing molecule can be modified to attach universal adapters (e.g., non-target nucleic acid sequences) to one or both ends of different target nucleic acid sequences. One or more universal primers attached to the target nucleic acids can provide sites for hybridization of the universal primers.The one or more universal primers bound to the target nucleic acid can be the same or different from each other.

[0047] As used herein, the terms "associated" or "associated with" have their conventional and ordinary meaning in the art in light of this specification. It can refer to two or more species that are identifiable as being co-located at some point in time. Association can refer to two or more species that are or were in similar containers. Association can refer to an informatic association. For example, digital information associated with two or more species can be stored and used to determine that one or more of these species were co-located at some point in time. Association can also refer to a physical association. In some embodiments, two or more associated species are "tethered," "bound," or "immobilized" to each other or to a common solid or semi-solid surface. Association can refer to covalent or non-covalent means for attaching a label to a solid or semi-solid support, such as a bead. Association can refer to a covalent bond between a target and a label. Association can include hybridization between two molecules (such as a target molecule and a label).

[0048] As used herein, the term "complementary" has its conventional and usual meaning in the art, taking this specification into consideration. It can refer to the ability for precise pairing between two nucleotides. For example, if a nucleotide at a given position in a nucleic acid is capable of hydrogen bonding with a nucleotide in another nucleic acid, the two nucleic acids are considered to be complementary to each other at that position. Complementarity between two single-stranded nucleic acid molecules can be "partial" when only a portion of the nucleotides bind, or it can be complete when there is overall complementarity between the single-stranded molecules. If a first nucleotide sequence is complementary to a second nucleotide sequence, the first nucleotide sequence can be said to be the "complement" of the second sequence. If a first nucleotide sequence is complementary to the reverse (i.e., the order of the nucleotides is reversed) sequence of the second sequence, the first nucleotide sequence can be said to be the "reverse complement" of the second sequence. As used herein, a "complementary" sequence can refer to the "complement" or "reverse complement" of a sequence. It is understood from this disclosure that when a molecule is capable of hybridizing to another molecule, it may be complementary or partially complementary to the hybridizing molecule.

[0049] As used herein, the term "digital counting" can refer to a method for estimating the number of target molecules in a sample. Digital counting can include determining the number of unique labels associated with targets in a sample. This method can be probabilistic in nature, transforming the problem of counting molecules from one of locating and identifying identical molecules to a series of digital present / absent problems involving the detection of a given set of labels.

[0050] As used herein, the term "label" or "labels" has its conventional and usual meaning in the art in light of this specification. They may refer to a nucleic acid code associated with a target in a sample. A label may, for example, comprise, consist essentially of, or consist of a nucleic acid label. A label may be an amplifiable label, entirely or partially. A label may be a sequenceable label, entirely or partially. A label may be a portion of a native nucleic acid that can be identified as unique. A label may comprise, consist essentially of, or consist of a known sequence. A label may include a nucleic acid sequence junction, e.g., a junction of a native and a non-native sequence. As used herein, the term "label" may be used synonymously with the terms "index," "tag," or "label tag." A label can convey information. For example, in various embodiments, a label may be used to determine the identity of a sample, the source of the sample, the identity of a cell, and / or a target.

[0051] As used herein, the term "non-depletion reservoir" can refer to a pool of barcodes (e.g., probabilistic barcodes) composed of many different labels. A non-depletion reservoir can contain many different barcodes so that when the non-depletion reservoir is associated with a pool of targets, each target is more likely to be associated with a unique barcode. The uniqueness of each labeled target molecule can be determined by the statistics of random selection and depends on the copy number of the same target molecule in the collection compared to the diversity of the labels. The size of the resulting set of labeled target molecules can be determined by the stochastic nature of the barcoding process, and analysis of the number of detected barcodes then allows for the calculation of the number of target molecules present in the original collection or sample. If the ratio of the copy number of the target molecule present to the number of unique barcodes is low, the labeled target molecule is highly unique (i.e., the probability that two or more target molecules are labeled with a given label is very low).

[0052] As used herein, the term "nucleic acid" has its conventional and usual meaning in the art in light of this specification. It refers to a polynucleotide sequence or a fragment thereof. A nucleic acid can comprise, consist essentially of, or consist of nucleotides. A nucleic acid can be exogenous or endogenous to a cell. A nucleic acid can be present in a cell-free environment. A nucleic acid can comprise, consist essentially of, or consist of a gene or a fragment thereof. A nucleic acid can comprise, consist essentially of, or consist of DNA. A nucleic acid can comprise, consist essentially of, or consist of RNA. A nucleic acid can comprise, consist essentially of, or consist of one or more analogs (e.g., modified backbones, sugars, or nucleobases). Some non-limiting examples of analogs include 5-bromouracil, peptide nucleic acids, xenonucleic acids, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein attached to a sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queosine, and wyosine. "Nucleic acid," "polynucleotide," "target polynucleotide," and "target nucleic acid" may be used interchangeably.

[0053] Nucleic acids can contain one or more modifications (e.g., base modifications, backbone modifications) to provide nucleic acids with new or improved characteristics (e.g., improved stability). Nucleic acids can contain a nucleic acid affinity tag. A nucleoside can be a base-sugar combination. The base portion of a nucleoside can be a heterocyclic base. The two most common classes of such heterocyclic bases are purines and pyrimidines. A nucleotide can be a nucleoside that further includes a phosphate group covalently linked to the sugar portion of the nucleoside. In nucleosides that include a pentofuranosyl sugar, the phosphate group can be attached to the 2', 3', or 5' hydroxyl moiety of the sugar. In forming nucleic acids, the phosphate group can covalently link adjacent nucleosides to one another to form a linear polymeric compound. Thus, the respective ends of this linear polymeric compound can be further linked to form a circular compound; however, linear compounds are generally preferred. Additionally, linear compounds may have internal nucleotide base complementarity and therefore may fold to produce fully or partially double-stranded compounds. Within nucleic acids, the phosphate groups may generally be referred to as forming the internucleoside backbone of the nucleic acid. The linkage or backbone may be a 3'-5' phosphodiester bond.

[0054] Nucleic acids can contain modified backbones and / or modified internucleoside linkages. Modified backbones can include those that retain a phosphorus atom in the backbone and those that do not have a phosphorus atom in the backbone. Suitable modified nucleic acid backbones containing a phosphorus atom therein can include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkylphosphotriesters, methyl and other alkyl phosphonates, such as 3'-alkylene phosphonates, 5'-alkylene phosphonates, chiral phosphonates, phosphinates, phosphoramidates including 3'-aminophosphoramidate and aminoalkylphosphoramidate, phosphorodiamidates, thienophosphoramidates, thienoalkylphosphonates, thienoalkylphosphotriesters, selenophosphates and boranophosphates having normal 3'-5' linkages, 2'-5' linkage analogs, and those having reverse polarity in which one or more internucleotide linkages are 3'-3', 5'-5', or 2'-2' linkages.

[0055] Nucleic acids can contain polynucleotide backbones formed by short chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short chain heteroatom or heterocyclic internucleoside linkages. These can include morpholino linkages (formed in part from the sugar portion of the nucleoside), siloxane backbones, sulfide, sulfoxide, and sulfone backbones, formacetyl and thioformacetyl backbones, methyleneformacetyl and thioformacetyl backbones, riboacetyl backbones, alkene-containing backbones, sulfamate backbones, methyleneimino and methylenehydrazino backbones, sulfonate and sulfonamide backbones, those with amide backbones, and others with mixed N, O, S, and CH2 moieties.

[0056] Nucleic acids can comprise, consist essentially of, or consist of nucleic acid mimetics. The term "mimetic" is intended to include polynucleotides in which only the furanose ring or both the furanose ring and the internucleotide linkage are replaced with non-furanose groups; replacement of only the furanose ring can also be referred to as a sugar surrogate. A heterocyclic base moiety or a modified heterocyclic base moiety can be retained for hybridization with an appropriate target nucleic acid. One such nucleic acid can be a peptide nucleic acid (PNA). In a PNA, the sugar backbone of a polynucleotide can be replaced with an amide-containing backbone, particularly an aminoethylglycine backbone. Nucleotides can be retained and directly or indirectly linked to aza nitrogen atoms in the amide portion of the backbone. The backbone in a PNA compound can contain two or more linked aminoethylglycine units, giving the PNA an amide-containing backbone. A heterocyclic base moiety can be directly or indirectly linked to an aza nitrogen atom in the amide portion of the backbone.

[0057] The nucleic acid may comprise, consist essentially of, or consist of a morpholino backbone structure. For example, the nucleic acid may contain a six-membered morpholino ring instead of a ribose ring. In some of these embodiments, phosphorodiamidate or other non-phosphodiester internucleoside linkages may replace the phosphodiester linkage.

[0058] Nucleic acids can comprise, consist essentially of, or consist of linked morpholino units (e.g., morpholino nucleic acids) having heterocyclic bases attached to morpholino rings. Linking groups can connect morpholino monomer units in morpholino nucleic acids. Nonionic morpholino-based oligomeric compounds can have fewer undesirable interactions with cellular proteins. Morpholino-based polynucleotides can be nonionic mimics of nucleic acids. Various compounds within the morpholino class can be linked using different linking groups. A further class of polynucleotide mimetics can be called cyclohexenyl nucleic acids (CeNA). The furanose ring normally present in nucleic acid molecules can be replaced with a cyclohexenyl ring. CeNA DMT-protected phosphoramidite monomers can be prepared and used for oligomeric compound synthesis using phosphoramidite chemistry. Incorporation of CeNA monomers into nucleic acid chains can increase the stability of DNA / RNA hybrids. CeNA oligoadenylates can form complexes with complementary nucleic acids with stability similar to that of native complexes. Further modifications include locked nucleic acids (LNAs), in which a 2'-hydroxyl group is attached to the 4' carbon atom of the sugar ring, forming a 2'-C,4'-C-oxymethylene linkage to form a bicyclic sugar moiety. The linkage can be a methylene (-CH2) group (where n is 1 or 2) bridging the 2' oxygen atom and the 4' carbon atom. LNAs and LNA analogs can exhibit very high duplex thermal stability with complementary nucleic acids (Tm = +3 to +10°C), stability against 3'-exonuclease degradation, and good solubility.

[0059] Nucleic acids may also include nucleobase (often simply referred to as "base") modifications or substitutions. As used herein, "unmodified" or "natural" nucleobases may include purine bases (e.g., adenine (A) and guanine (G)) and pyrimidine bases (e.g., thymine (T), cytosine (C), and uracil (U)). Modified nucleobases include other synthetic and natural nucleobases, such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C=C-CH3) uracil and other alkynyl derivatives of cytosine and pyrimidine bases, 6-azouracil, cytosine ... These may include tosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo, particularly 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-aminoadenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine and 3-deazaguanine and 3-deazaadenine.Modified nucleobases include tricyclic pyrimidines, such as phenoxazine cytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), G-clamps such as substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), G-clamps such as substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), carbazole cytidines (2H-pyrimido(4,5-b)indol-2-one), and pyridoindole cytidines (H-pyrido(3',2':4,5)pyrrolo[2,3-d]pyrimidin-2-one).

[0060] As used herein, the term "sample" can refer to a composition containing a target. Samples suitable for analysis by the disclosed methods, devices, and systems include cells, tissues, organs, or organisms. In some embodiments, a sample comprises, consists essentially of, or consists of a single cell. In some embodiments, a sample comprises, consists essentially of, or consists of at least 100,000, 200,000, 300,000, 500,000, 800,000, or 1,000,000 single cells.

[0061] As used herein, the terms "sampling device" or "device" may refer to a device capable of taking a section of a sample and / or depositing a section on a substrate. A sample device may refer to, for example, a fluorescence activated cell sorting (FACS) machine, a cell sorter machine, a biopsy needle, a biopsy device, a tissue sectioning device, a microfluidic device, a blade grid, and / or a microtome.

[0062] As used herein, the term "solid support" has its conventional and ordinary meaning in the art in light of this specification. It can refer to a discrete solid or semi-solid surface to which multiple barcodes (e.g., stochastic barcodes) can be attached. A solid support can include any type of solid, porous, or hollow sphere, ball, bearing, cylinder, or other similar structure composed of plastic, ceramic, metal, or polymeric material (e.g., hydrogel) to which nucleic acids can be immobilized (e.g., covalently or non-covalently). A solid support can include discrete particles that can be spherical (e.g., microspheres) or have non-spherical or irregular shapes, such as cubic, rectangular, pyramidal, cylindrical, conical, ellipsoidal, or disc-shaped. Beads can be non-spherical in shape. A plurality of solid supports spaced apart in an array can also be substrate-free. A solid support can be used synonymously with the term "bead." For any embodiment herein in which a barcode is immobilized on a solid support, particle, bead, or the like, it is contemplated that the barcode may also be separated into droplets (e.g., microdroplets), such as, for example, hydrogel droplets, or into wells of a substrate, such as microwells, or into chambers of a fluidic device (e.g., a microfluidic device). Thus, wherever sorting, sorting, or distribution of nucleic acids by "solid support" (e.g., beads) is described herein, distribution into a fluid (e.g., droplets, such as microdroplets) or physical space, such as a microwell (e.g., on a multiwell plate) or chamber (e.g., in a fluidic device) is also expressly contemplated.

[0063] As used herein, the term "probabilistic barcode" may refer to a polynucleotide sequence comprising a label of the present disclosure. A probabilistic barcode may be a polynucleotide sequence that can be used for probabilistic barcoding. A probabilistic barcode may be used to quantify a target in a sample. A probabilistic barcode may be used to control errors that may occur after associating a label with a target. For example, a probabilistic barcode may be used to evaluate amplification or sequencing errors. A probabilistic barcode associated with a target may be referred to as a probabilistic barcode target or a probabilistic barcode tag target.

[0064] As used herein, the term "gene-specific probabilistic barcode" may refer to a polynucleotide sequence that includes a label and a gene-specific target binding region. A probabilistic barcode may be a polynucleotide sequence that can be used for probabilistic barcoding. A probabilistic barcode may be used to quantify a target in a sample. A probabilistic barcode may be used to control errors that may occur after associating a label with a target. For example, a probabilistic barcode may be used to evaluate amplification or sequencing errors. A probabilistic barcode associated with a target may be referred to as a probabilistic barcode target or a probabilistic barcode tag target.

[0065] As used herein, the term "probabilistic barcoding" can refer to random labeling (e.g., barcoding) of nucleic acids. Probabilistic barcoding can use a recursive Poisson strategy to associate a label with a target and quantify the label associated with the target. As used herein, the term "probabilistic barcoding" can be used synonymously with "probabilistic labeling."

[0066] As used herein, the term "target" has its conventional and ordinary meaning in the art in light of this specification. It may refer to a composition that can be associated with a barcode (e.g., a probabilistic barcode). Exemplary targets suitable for analysis by the disclosed methods, devices, and systems include oligonucleotides, DNA, RNA, mRNA, microRNA, tRNA, and the like. A target may be single-stranded or double-stranded. In some embodiments, a target may be a protein, peptide, or polypeptide. In some embodiments, a target is a lipid. As used herein, "target" may be used synonymously with "species."

[0067] As used herein, the term "reverse transcriptase" has its conventional and ordinary meaning in the art in light of this specification. It can refer to a group of enzymes that have reverse transcriptase activity (i.e., that catalyze the synthesis of DNA from an RNA template). Generally, such enzymes include, but are not limited to, retroviral reverse transcriptases, retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retron reverse transcriptases, bacterial reverse transcriptases, group II intron-derived reverse transcriptases, and mutants, variants, or derivatives thereof. Non-retroviral reverse transcriptases include non-LTR retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retron reverse transcriptases, and group II intron reverse transcriptases. Examples of group II intron reverse transcriptases include the Lactococcus lactis LI.LtrB intron reverse transcriptase, the Thermosynechococcus elongatus TeI4c intron reverse transcriptase, or the Geobacillus stearothermophilus GsI-IIC intron reverse transcriptase. Other classes of reverse transcriptases can include many classes of non-retroviral reverse transcriptases (i.e., retrons, group II introns, and particularly diversity-generating retroelements).

[0068] The terms "universal adapter primer," "universal primer adapter," or "universal adapter sequence" are used interchangeably to refer to a nucleotide sequence that can be used to hybridize a barcode (e.g., a stochastic barcode) to generate a gene-specific barcode. The universal adapter sequence can be, for example, a known sequence that is universal for all barcodes used in the methods of the present disclosure. For example, when multiple targets are labeled using the methods disclosed herein, each of the target-specific sequences can be attached to the same universal adapter sequence. In some embodiments, two or more universal adapter sequences can be used in the methods disclosed herein. For example, when multiple targets are labeled using the methods disclosed herein, at least two of the target-specific sequences are attached to different universal adapter sequences. The universal adapter primer and its complement can be included in two oligonucleotides, one of which contains the target-specific sequence and the other of which contains the barcode. For example, the universal adapter sequence can be part of an oligonucleotide that contains a target-specific sequence to generate a nucleotide sequence complementary to the target nucleic acid. A second oligonucleotide comprising the complementary sequence of the barcode and universal adapter sequence can hybridize with the nucleotide sequence to generate a target-specific barcode (e.g., a target-specific stochastic barcode). In some embodiments, the universal adapter primer has a different sequence than the universal PCR primer used in the disclosed methods.

[0069] Barcode Barcoding, such as probabilistic barcoding, is described, for example, in U.S. Patent Application Publication No. 2015 / 0299784, WO 2015 / 031691, and Fu et al., Proc Natl Acad Sci USA 2011 May 31;108(22):9026-31 (the contents of each of these publications are incorporated herein by reference in their entirety). In some embodiments, the barcodes disclosed herein can be probabilistic barcodes, which can be polynucleotide sequences that can be used to probabilistically label (e.g., barcode, tag) targets. A barcode may be referred to as a probabilistic barcode if the ratio of the number of distinct barcode sequences in the probabilistic barcode to the number of occurrences of any of the labeled targets can be 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or a number or range between any two of these values ​​or an approximation of such a value. The targets may be mRNA species that include mRNA molecules with identical or nearly identical sequences. A barcode may be referred to as a probabilistic barcode if the ratio of the number of distinct barcode sequences of the probabilistic barcode to the number of occurrences of any of the labeled targets is at least or at most 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100:1. The barcode sequences of a probabilistic barcode may be referred to as molecular labels.

[0070] A barcode, e.g., a probabilistic barcode, can include one or more labels. Exemplary labels can include a universal label, a cell label, a barcode sequence (e.g., a molecular label), a sample label, a plate label, a spatial label, and / or a pre-spatial label. FIG. 1 shows an exemplary barcode 104 having a spatial label. The barcode 104 can include a 5' amine that can link the barcode to a solid support 105. The barcode can include a universal label, a dimensional label, a spatial label, a cell label, and / or a molecular label. The barcode can include a universal label, a cell label, and a molecular label. The barcode can include a universal label, a spatial label, a cell label, and a molecular label. The barcode can include a universal label, a dimensional label, a cell label, and a molecular label. The order of various labels (including, but not limited to, a universal label, a dimensional label, a spatial label, a cell label, and / or a molecular label) in a barcode can vary. For example, as shown in FIG. 1, the universal label can be the 5'-most label and the molecular label can be the 3'-most label. The spatial label, dimensional label, and cellular label can be in any order. In some embodiments, the universal label, spatial label, dimensional label, cellular label, and molecular label can be in any order. The barcode can include a target binding region. The target binding region can interact with a target in a sample (e.g., a target nucleic acid, RNA, mRNA, DNA). For example, the target binding region can include an oligo(dT) sequence that can interact with the poly(A) tail of an mRNA. In some cases, the labels of the barcode (e.g., the universal label, dimensional label, spatial label, cellular label, and barcode sequence) can be separated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides.

[0071] Labels, e.g., cellular labels, can include a unique set of nucleic acid subsequences of a defined length, e.g., seven nucleotides each (corresponding to the number of bits used in some Hamming error-correcting codes), which can be designed to confer error-correction capabilities. The set of error-correcting subsequences includes seven nucleotide sequences, which can be designed so that any pairwise combination of sequences in the set exhibits a defined "genetic distance" (or number of mismatched bases); for example, a set of error-correcting subsequences can be designed to exhibit a genetic distance of three nucleotides. In this case, review of the error-correcting sequences within a set of sequence data for a labeled target nucleic acid molecule (described in more detail below) allows for the detection or correction of amplification or sequencing errors. In some embodiments, the length of the nucleic acid subsequences used to create the error-correcting code can vary; e.g., they can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 31, 40, 50 nucleotides in length, or a number or range between or approximating any two of these values. In some embodiments, nucleic acid subsequences of other lengths can be used to create error correcting codes.

[0072] The barcode may include a target binding region. The target binding region may interact with a target in a sample. The target may be or include ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNA containing a poly(A) tail, or any combination thereof. In some embodiments, the multiple targets may include deoxyribonucleic acid (DNA).

[0073] In some embodiments, the target binding region may include an oligo(dT) sequence that can interact with the poly(A) tail of mRNA. One or more of the labels of the barcode (e.g., universal label, dimensional label, spatial label, cellular label, and barcode sequence (e.g., molecular label)) may be separated from one or two of the remaining labels of the barcode by a spacer. The spacer may be, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides. In some embodiments, none of the labels of the barcode are separated by a spacer.

[0074] Universal Signage A barcode may include one or more universal labels. In some embodiments, the one or more universal labels may be the same for all barcodes in a set of barcodes bound to a given solid support. In some embodiments, the one or more universal labels may be the same for all barcodes bound to a plurality of beads. In some embodiments, the universal label may include a nucleic acid sequence capable of hybridizing to a sequencing primer. The sequencing primer may be used to sequence the barcodes comprising the universal label. The sequencing primer (e.g., a universal sequencing primer) may include a sequencing primer associated with a high-throughput sequencing platform. In some embodiments, the universal label may include a nucleic acid sequence capable of hybridizing to a PCR primer. In some embodiments, the universal label may include a nucleic acid sequence capable of hybridizing to a sequencing primer and a PCR primer. The nucleic acid sequence of a universal label capable of hybridizing to a sequencing primer or a PCR primer may be referred to as a primer binding site. The universal label may include a sequence that can be used to initiate transcription of the barcode. The universal label may include a sequence that can be used to extend the barcode or a region within the barcode. A universal label can be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range between or approximating any two of these values. For example, a universal label can comprise at least about 10 nucleotides. A universal label can be at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length. In some embodiments, a cleavable linker or modified nucleotide can be part of the universal label sequence that allows for cleavage and removal of the barcode from the support.

[0075] dimensional indicator A barcode can include one or more dimensional labels. In some embodiments, a dimensional label can include a nucleic acid sequence that provides information about the dimension in which the labeling (e.g., stochastic labeling) occurred. For example, a dimensional label can provide information about the time at which a target was barcoded. A dimensional label can be associated with the time of sample barcoding (e.g., stochastic barcoding). A dimensional label can be activated at the time of labeling. Different dimensional labels can be activated at different time points. A dimensional label provides information about the order in which targets, groups of targets, and / or samples were barcoded. For example, a cell population can be barcoded in the G0 phase of the cell cycle. Cells can be pulsed again with a barcode (e.g., a stochastic barcode) in the G1 phase of the cell cycle. Cells can be pulsed again with a barcode in the S phase of the cell cycle, and so on. The barcode for each pulse (e.g., each phase of the cell cycle) can include a different dimensional label. In this way, the dimensional label provides information about which phase of the cell cycle the target was labeled in. Dimensional labels can probe many different biological time periods. Exemplary biological time periods can include, but are not limited to, cell cycle, transcription (e.g., transcription initiation), and transcript degradation. In another example, a sample (e.g., a cell, a cell population) can be labeled before and / or after drug treatment and / or therapy. Changes in the copy number of unique targets can be an indicator of the sample's response to the drug and / or therapy.

[0076] Dimensional labels may be activatable. Activatable dimensional labels may be activated at specific times. Activatable labels may, for example, be constitutively activated (e.g., do not switch off). Activatable dimensional labels may, for example, be reversibly activated (e.g., they can be switched on and off). Dimensional labels may, for example, be reversibly activatable at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more times. Dimensional labels may, for example, be reversibly activatable at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more times. In some embodiments, dimensional labels may be activated by fluorescence, light, a chemical event (e.g., cleavage, ligation of another molecule, addition of a modification (e.g., pegylation, sumoylation, acetylation, methylation, deacetylation, demethylation), a photochemical event (e.g., photocaging), and the introduction of a non-natural nucleotide.

[0077] In some embodiments, the dimension labels may be the same for all barcodes (e.g., stochastic barcodes) attached to a given solid support (e.g., a bead), but may be different for different solid supports (e.g., beads). In some embodiments, at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100% of the barcodes on the same solid support may comprise the same dimension label. In some embodiments, at least 60% of the barcodes on the same solid support may comprise the same dimension label. In some embodiments, at least 95% of the barcodes on the same solid support may comprise the same dimension label.

[0078] For multiple solid supports (e.g., beads), 10 6There may be about or more unique dimension label sequences. Dimension labels can be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or any number or range between or approximations of any two of these values. Dimension labels can be at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length. Dimension labels can comprise from about 5 to about 200 nucleotides. Dimension labels can comprise from about 10 to about 150 nucleotides. Dimension labels can comprise from about 20 to about 125 nucleotides in length.

[0079] spatial sign A barcode may include one or more spatial labels. In some embodiments, a spatial label may include a nucleic acid sequence that provides information about the spatial orientation of a target molecule associated with the barcode. A spatial label may be associated with a coordinate in a sample. The coordinate may be a fixed coordinate. For example, the coordinate may be fixed relative to a substrate. The spatial label may be referenced to a two-dimensional or three-dimensional grid. The coordinate may be fixed relative to a landmark. A landmark may be identifiable in space. A landmark may be an imageable structure. A landmark may be a biological structure, e.g., an anatomical landmark. A landmark may be a cellular landmark, e.g., an organelle. A landmark may be a non-natural landmark, such as a color code, a barcode, a structure with an identifiable identifier, such as magnetic, fluorescent, radioactive, or a unique size or shape. A spatial label may be associated with a physical compartment (e.g., a well, a container, or a droplet). In some embodiments, multiple spatial labels are used together to code one or more locations in space.

[0080] Spatial labels can be the same for all barcodes attached to a given solid support (e.g., a bead), but can be different for different solid supports (e.g., beads). In some embodiments, the percentage of barcodes on the same solid support that contain the same spatial label can be 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values, or an approximation of such a value. In some embodiments, the percentage of barcodes on the same solid support that contain the same spatial label can be at least or at most 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. In some embodiments, at least 60% of the barcodes on the same solid support can contain the same spatial label. In some embodiments, at least 95% of the barcodes on the same solid support can contain the same spatial label.

[0081] For multiple solid supports (e.g., beads), 10 6 There may be about or more unique spatial marker sequences. Spatial markers may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or any number or range between or approximations of any two of these values. Spatial markers may be at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length. Spatial markers may comprise about 5 to about 200 nucleotides. Spatial markers may comprise about 10 to about 150 nucleotides. Spatial markers may comprise about 20 to about 125 nucleotides in length.

[0082] cell labeling A barcode (e.g., a probabilistic barcode) may include one or more cell labels. In some embodiments, the cell label may include a nucleic acid sequence that provides information for determining which target nucleic acid originates from which cell. In some embodiments, the cell label is the same for all barcodes attached to a given solid support (e.g., a bead) but different for different solid supports (e.g., beads). In some embodiments, the percentage of barcodes on the same solid support that contain the same cell label may be 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values, or an approximation of such a value. In some embodiments, the percentage of barcodes on the same solid support that contain the same cell label may be 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%, or an approximation of such a value. For example, at least 60% of barcodes on the same solid support may contain the same cell label. As another example, at least 95% of the barcodes on the same solid support may contain the same cell label.

[0083] For multiple solid supports (e.g., beads), 10 6 There may be about or more unique cellular marker sequences. A cellular marker can be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range between any two of these values, or an approximation of such a value. A cellular marker can be at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length. For example, a cellular marker can comprise about 5 to about 200 nucleotides. As another example, a cellular marker can comprise about 10 to about 150 nucleotides. As yet another example, a cellular marker can comprise about 20 to about 125 nucleotides in length.

[0084] Barcode sequence The barcode may include one or more barcode sequences. In some embodiments, the barcode sequence may include a nucleic acid sequence that provides information for identifying the specific type of target nucleic acid species hybridized to the barcode. The barcode sequence may include a nucleic acid sequence that provides a counter (e.g., provides an approximation) for the specific presence of the target nucleic acid species hybridized to the barcode (e.g., target binding region).

[0085] In some embodiments, a diverse set of barcode sequences is attached to a given solid support (e.g., a bead). 2 , 10 3 、 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 Or there may be a number or range between any two of these values ​​or an approximation of such values ​​of unique molecular label sequences. For example, the plurality of barcodes may include about 6561 barcode sequences with unique sequences. As another example, the plurality of barcodes may include about 65536 barcode sequences with unique sequences. In some embodiments, there may be at least or at most 10 2 , 10 3 、 10 4 , 10 5 , 10 6 , 10 7 , 10 8 or 10 9 There may be a unique barcode sequence of 1000. The unique molecular tag sequence may be attached to a given solid support (e.g., a bead). In some embodiments, the unique molecular tag sequence is partially or entirely encompassed by a particle (e.g., a hydrogel bead).

[0086] The length of the barcode may vary in different implementations. For example, the barcode may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range between any two of these values, or an approximation of such a value. As another example, the barcode may be at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length.

[0087] molecular label A barcode (e.g., a probabilistic barcode) can include one or more molecular labels. A molecular label can include a barcode sequence. In some embodiments, a molecular label can include a nucleic acid sequence that provides information for identifying the specific type of target nucleic acid species hybridized to the barcode. A molecular label can include a nucleic acid sequence that provides a counter for the specific presence of a target nucleic acid species hybridized to the barcode (e.g., a target binding region).

[0088] In some embodiments, a diverse set of molecular labels is attached to a given solid support (e.g., a bead). 2 , 10 3 、 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 Or there may be a number or range between any two of these values ​​or an approximation of such values. For example, the plurality of barcodes may include about 6561 molecular labels with unique sequences. As another example, the plurality of barcodes may include about 65536 molecular labels with unique sequences. In some embodiments, at least or at most 10 2 , 10 3 、 10 4 , 10 5 , 10 6 , 10 7 , 10 8 or 109 There can be a unique molecular tag sequence of 100. A barcode having a unique molecular tag sequence can be attached to a given solid support (e.g., a bead).

[0089] In barcoding using multiple probabilistic barcodes (e.g., probabilistic barcoding), the ratio of the number of distinct molecular label sequences to the number of occurrences of any of the targets can be 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or a number or range between any two of these values ​​or an approximation of such a value. The targets can be mRNA species that include mRNA molecules with identical or nearly identical sequences. In some embodiments, the ratio of the number of different molecular label sequences to the number of occurrences of any of the targets is at least or at most 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100:1.

[0090] A molecular label can be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range between or approximation of any two of these values. A molecular label can be at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length.

[0091] Target binding region The barcode may include one or more target binding regions, such as a capture probe. In some embodiments, the target binding region may hybridize with a target of interest. In some embodiments, the target binding region may include a nucleic acid sequence that specifically hybridizes to a target (e.g., a target nucleic acid, a target molecule, e.g., a cellular nucleic acid to be analyzed), such as a specific gene sequence. In some embodiments, the target binding region may include a nucleic acid sequence that can bind (e.g., hybridize) to a specific position of a specific target nucleic acid. In some embodiments, the target binding region may include a nucleic acid sequence capable of specific hybridization to a restriction enzyme site overhang (e.g., an EcoRI sticky end overhang). The barcode may then be ligated to any nucleic acid molecule that contains a sequence complementary to the restriction site overhang.

[0092] In some embodiments, the target binding region can include a non-specific target nucleic acid sequence. A non-specific target nucleic acid sequence can refer to a sequence that can bind to multiple target nucleic acids regardless of the specific sequence of the target nucleic acid. For example, the target binding region can include a random multimer sequence or an oligo(dT) sequence that hybridizes to the poly(A) tail of an mRNA molecule. The random multimer sequence can be, for example, a random dimer, trimer, quatramer, pentamer, hexamer, septamer, octamer, nonamer, decamer, or higher-order multimer sequence of any length. In some embodiments, the target binding region is the same for all barcodes bound to a given bead. In some embodiments, the target binding regions of multiple barcodes bound to a given bead can include two or more different target binding sequences. The target binding region can be 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range between or approximating any two of these values. The target binding region can be at most about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 or more nucleotides in length.

[0093] In some embodiments, the target binding region can include an oligo(dT) that can hybridize to an mRNA containing a polyadenylated end. The target binding region can be gene-specific. For example, the target binding region can be configured to hybridize to a specific region of the target. In some embodiments, the target binding region does not include an oligo(dT). The target binding region can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nucleotides in length, or a number or range between any two of these values, or an approximation of such a value. The target binding region can be at least or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. The target binding region can be about 5-30 nucleotides in length. When a barcode includes a gene-specific target binding region, the barcode may be referred to herein as a gene-specific barcode.

[0094] Orientation A probabilistic barcode (e.g., a probabilistic barcode) can include one or more orientations that can be used to orient (e.g., align) the barcode. The barcode can include a moiety for isoelectric focusing. Different barcodes can include different isoelectric focusing points. When these barcodes are introduced into a sample, the sample can undergo isoelectric focusing to orient the barcodes into a known configuration. In this manner, orientations can be used to create a known map of barcodes in the sample. Exemplary orientations include electrophoretic mobility (e.g., based on the size of the barcode), isoelectric point, spin, conductivity, and / or self-assembly. For example, a barcode with orientations for self-assembly can self-assemble into a specific orientation upon activation (e.g., a nucleic acid nanostructure).

[0095] affinity A barcode (e.g., a probabilistic barcode) may include one or more affinities. For example, a spatial label may include an affinity. The affinity may include a chemical and / or biological moiety that can facilitate binding of the barcode to another entity (e.g., a cellular receptor). For example, the affinity may include an antibody, e.g., an antibody specific to a particular moiety (e.g., a receptor) on a sample. In some embodiments, the antibody may direct the barcode to a particular cell type or molecule. A particular cell type or molecule and / or a target in its vicinity may be labeled (e.g., stochastically labeled). Because the antibody can direct the barcode to a specific location, the affinity, in some embodiments, can provide spatial information in addition to the nucleotide sequence of the spatial label. The antibody may be a therapeutic antibody, e.g., a monoclonal or polyclonal antibody. The antibody may be humanized or chimeric. The antibody may be a naked antibody or a fusion antibody.

[0096] Antibodies can be full-length (i.e., naturally occurring or generated by conventional immunoglobulin gene fragment recombination processes) immunoglobulin molecules (e.g., IgG antibodies) or immunoreactive (i.e., specific binding) portions of immunoglobulin molecules, such as antibody fragments.

[0097] An antibody fragment can be a portion of an antibody, such as, for example, F(ab')2, Fab', Fab, Fv, or sFv. In some embodiments, an antibody fragment can bind to the same antigen recognized by a full-length antibody. An antibody fragment can include an isolated fragment consisting of the variable region of an antibody, such as an "Fv" fragment consisting of the variable regions of the heavy and light chains, and a recombinant single-chain polypeptide molecule in which the variable regions of the light and heavy chains are connected by a peptide linker ("scFv protein"). Exemplary antibodies can include, but are not limited to, antibodies against cancer cells, antibodies against viruses, antibodies that bind to cell surface receptors (CD8, CD34, CD45), and therapeutic antibodies.

[0098] Universal Adapter Primer A barcode can include one or more universal adapter primers. For example, a gene-specific barcode, such as a gene-specific probability barcode, can include a universal adapter primer. A universal adapter primer can refer to a nucleotide sequence that is universal for all barcodes. A universal adapter primer can be used to construct a gene-specific barcode. A universal adapter primer can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nucleotides in length, or a number or range between any two of these nucleotide lengths, or an approximation of such a value. The universal adapter primer can be at least or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. The universal adapter primer can be 5 to 30 nucleotides in length.

[0099] Linker When a barcode includes two or more types of labels (e.g., two or more cellular labels or two or more barcode sequences, e.g., one molecular label), the labels may incorporate a linker label sequence. The linker label sequence may be at least about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or more nucleotides in length. The linker label sequence may be at most about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or more nucleotides in length. In some cases, the linker label sequence is 12 nucleotides in length. The linker label sequence may be used to facilitate synthesis of the barcode. The linker label may include an error-correcting (e.g., Hamming) code.

[0100] solid support In some embodiments, barcodes, such as the stochastic barcodes disclosed herein, can be associated with a solid support. The solid support can be, for example, a synthetic particle. In some embodiments, some or all of the barcode sequences (e.g., first barcode sequences), such as molecular labels of stochastic barcodes, of a plurality of barcodes (e.g., a first plurality of barcodes) on a solid support differ by at least one nucleotide. The cellular labels of barcodes on the same solid support can be the same. The cellular labels of barcodes on different solid supports can differ by at least one nucleotide. For example, a first cellular label of a first plurality of barcodes on a first solid support can have the same sequence, and a second cellular label of a second plurality of barcodes on a second solid support can have the same sequence. The first cellular label of a first plurality of barcodes on a first solid support and the second cellular label of a second plurality of barcodes on a second solid support can differ by at least one nucleotide. The cellular labels can be, for example, about 5-20 nucleotides in length. The barcode sequences can be, for example, about 5-20 nucleotides in length. The synthetic particles can be, for example, beads.

[0101] The beads can be, for example, silica gel beads, controlled pore glass beads, magnetic beads, Dynabeads, Sephadex / Sepharose beads, cellulose beads, polystyrene beads, or any combination thereof. The beads can comprise materials such as polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogels, paramagnetic materials, ceramics, plastics, glass, methylstyrene, acrylic polymers, titanium, latex, Sepharose, cellulose, nylon, silicone, or any combination thereof.

[0102] In some embodiments, the beads may be polymer beads, such as deformable beads or gel beads, which are functionalized with barcodes or probabilistic barcodes (e.g., gel beads from 10X Genomics, San Francisco, CA). In some implementations, the gel beads may comprise a polymer-based gel. Gel beads may be made, for example, by encapsulating one or more polymer precursors in droplets. Gel beads may be made by exposing the polymer precursors to an accelerant (e.g., tetramethylethylenediamine (TEMED)).

[0103] In some embodiments, the particles can be disintegrable (e.g., dissolvable or degradable). For example, polymer beads can dissolve, melt, or decompose under desired conditions. The desired conditions can include environmental conditions. The desired conditions can result in the dissolution, melting, or decomposition of the polymer beads in a controlled manner. Gel beads can dissolve, melt, or decompose due to a chemical stimulus, a physical stimulus, a biological stimulus, a thermal stimulus, a magnetic stimulus, an electrical stimulus, an optical stimulus, or any combination thereof.

[0104] For example, analytes and / or reagents, such as oligonucleotide barcodes, can be coupled / immobilized to the interior surface of gel beads (e.g., the interior accessible by diffusion of the oligonucleotide barcodes and / or the material used to create the oligonucleotide barcodes) and / or to the exterior surface of gel beads or any other microcapsules described herein. Coupling / immobilization can be via any form of chemical bond (e.g., covalent bond, ionic bond) or physical phenomenon (e.g., van der Waals forces, dipole-dipole interactions, etc.). In some embodiments, the coupling / immobilization of reagents to gel beads or any other microcapsules described herein can be reversible, such as, for example, via a labile moiety (e.g., a chemical crosslinker, including those described herein). Upon application of a stimulus, the labile moiety can be cleaved, releasing the immobilized reagent. In some embodiments, the labile moiety is a disulfide bond. For example, if an oligonucleotide barcode is immobilized to a gel bead via a disulfide bond, exposure of the disulfide bond to a reducing agent can cleave the disulfide bond and release the oligonucleotide barcode from the bead. The labile moiety may be included as part of a gel bead or microcapsule, as part of a chemical linker connecting a reagent or analyte to the gel bead or microcapsule, and / or as part of the reagent or analyte. In some embodiments, at least one barcode of the plurality of barcodes may be immobilized on a particle, partially immobilized on a particle, encapsulated in a particle, partially encapsulated in a particle, or any combination thereof.

[0105] In some embodiments, the gel beads may comprise a variety of different polymers, including but not limited to polymers, thermosensitive polymers, photosensitive polymers, magnetic polymers, pH-sensitive polymers, salt-sensitive polymers, chemically sensitive polymers, polyelectrolytes, polysaccharides, peptides, proteins, and / or plastics. The polymer may include, but is not limited to, materials such as poly(N-isopropylacrylamide) (PNIPAAm), poly(styrene sulfonate) (PSS), poly(allylamine) (PAAm), poly(acrylic acid) (PAA), poly(ethyleneimine) (PEI), poly(diallyldimethyl-ammonium chloride) (PDADMAC), poly(pyrrole) (PPy), poly(vinylpyrrolidone) (PVPON), poly(vinylpyridine) (PVP), poly(methacrylic acid) (PMAA), poly(methyl methacrylate) (PMMA), polystyrene (PS), poly(tetrahydrofuran) (PTHF), poly(phthalaldehyde) (PTHF), poly(hexylviologen) (PHV), poly(L-lysine) (PLL), poly(L-arginine) (PARG), poly(lactic-co-glycolic acid) (PLGA).

[0106] Many chemical stimuli can be used to cause bead rupture, dissolution, or degradation. Examples of these chemical changes can include, but are not limited to, pH-mediated changes to the bead wall, bead wall collapse via chemical scission of cross-links, triggering bead wall depolymerization, and bead wall switching reactions. Bulk changes can also be used to trigger bead rupture.

[0107] Bulk or physical changes to microcapsules via various stimuli also offer many advantages in designing capsules to release reagents. Bulk or physical changes occur on a macroscopic scale, where bead rupture is the result of mechanical physical forces induced by the stimulus. These processes can include, but are not limited to, pressure-induced rupture, bead wall melting, or bead wall porosity changes.

[0108] Biological stimuli can also be used to trigger the disruption, dissolution, or degradation of beads. Generally, biological triggers are similar to chemical triggers, but many examples use biomolecules or molecules commonly present in biological systems, such as enzymes, peptides, sugars, fatty acids, and nucleic acids. For example, beads may contain polymers with peptide crosslinks that are susceptible to cleavage by specific proteases. More specifically, one example may contain microcapsules containing GFLGK peptide crosslinks. Addition of a biological trigger, such as the protease cathepsin B, cleaves the shell wall's peptide crosslinks, releasing the contents of the beads. In other cases, the protease may be heat-activated. In another example, beads contain a shell wall containing cellulose. Addition of the hydrolytic enzyme chitosan serves as a biological trigger for cleavage of the cellulose bonds, depolymerizing the shell wall, and releasing its internal contents.

[0109] Beads can also be induced to release their contents upon application of a thermal stimulus. A change in temperature can cause various changes in the beads. A change in heat can cause the beads to melt, causing the bead walls to collapse. In other cases, heat can increase the internal pressure of the beads' internal components, causing the beads to break or burst. In still other cases, heat can transform the beads into a shrunken, dehydrated state. Heat can also act on the thermosensitive polymers within the bead walls, causing the beads to break.

[0110] The inclusion of magnetic nanoparticles in the bead walls of microcapsules can trigger bead rupture as well as guide multiple beads. The devices of the present disclosure can include magnetic beads for any purpose. In one example, the incorporation of Fe3O4 nanoparticles into polyelectrolyte-containing beads triggers rupture in the presence of an oscillating magnetic field stimulus.

[0111] Beads can also be disrupted, dissolved, or decomposed as a result of electrical stimulation. Similar to the magnetic particles described in the previous section, electrically sensitive beads can also trigger bead rupture as well as other functions such as alignment under an electric field, conductivity, or redox reactions. In one example, beads containing electrically sensitive materials are aligned under an electric field so that the release of internal reagents can be controlled. In another example, an electric field can induce redox reactions within the bead wall itself, thereby increasing porosity.

[0112] Optical stimulation can also be used to disrupt the beads. Numerous optical triggers are possible, including systems using various molecules such as nanoparticles and chromophores that can absorb photons of specific wavelengths. For example, metal oxide coatings can be used as capsule triggers. UV irradiation of SiO2-coated polyelectrolyte capsules can result in the collapse of the bead wall. In yet another example, photoswitch materials such as azobenzene groups can be incorporated into the bead wall. Upon application of UV or visible light, chemicals such as these undergo reversible cis-trans isomerization upon absorption of a photon. In this embodiment, the incorporation of a photonic switch can cause the bead wall to collapse or become more porous upon application of the optical trigger.

[0113] For example, in a non-limiting example of barcoding (e.g., probabilistic barcoding) shown in FIG. 2, after introducing cells, such as single cells, into multiple microwells of a microwell array in block 208, beads may be introduced into multiple microwells of the microwell array in block 212. Each microwell may contain one bead. The beads may contain multiple barcodes. The barcodes may include 5' amine regions attached to the beads. The barcodes may include a universal label, a barcode sequence (e.g., a molecular label), a target binding region, or any combination thereof.

[0114] The barcodes disclosed herein can be associated with (e.g., attached to) solid supports (e.g., beads). The barcodes attached to the solid supports can include barcode sequences selected from a group including at least 100 or 1000 barcode sequences, each having a unique sequence. In some embodiments, different barcodes attached to the solid supports can include barcodes with different sequences. In some embodiments, a certain percentage of the barcodes attached to the solid supports include the same cell label. For example, the percentage can be 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values, or an approximation of such a value. As another example, the percentage can be at least or at most 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. In some embodiments, the barcodes attached to the solid supports can have the same cell label. The barcodes attached to the different solid supports can have different cell markers selected from a group comprising at least 100 or 1000 cell markers with unique sequences.

[0115] The barcodes disclosed herein can be associated with (e.g., attached to) a solid support (e.g., a bead). In some embodiments, barcoding multiple targets in a sample can be performed using a solid support comprising multiple synthetic particles associated with multiple barcodes. In some embodiments, the solid support can comprise multiple synthetic particles associated with multiple barcodes. The spatial labeling of multiple barcodes on different solid supports can differ by at least one nucleotide. The solid support can comprise multiple barcodes, for example, in two or three dimensions. The synthetic particles can be beads. The beads can be silica gel beads, controlled pore glass beads, magnetic beads, Dynabeads, Sephadex / Sepharose beads, cellulose beads, polystyrene beads, or any combination thereof. The solid support can comprise a polymer, matrix, hydrogel, needle array device, antibody, or any combination thereof. In some embodiments, the solid support can be free-floating. In some embodiments, the solid support can be embedded in a semi-solid or solid array. The barcodes need not be attached to the solid support. The barcodes can be individual nucleotides. The barcode can be associated with a substrate. In some embodiments, the barcode can be associated with a single cell in a compartment, such as a droplet, such as a microdroplet, or a well, such as a microwell, of a substrate (e.g., on a multiwell plate) or chamber (e.g., in a fluidic device). An example of a droplet can include a hydrogel droplet. The barcodes in the compartment can be immobilized on a solid support or they can be free in solution.

[0116] As used herein, the terms "tethered," "attached," and "immobilized" are used interchangeably and can refer to covalent or non-covalent means for attaching a barcode to a solid support. Any of a variety of different solid supports can be used as solid supports to attach pre-synthesized barcodes or for in situ solid phase synthesis of barcodes.

[0117] In some embodiments, the solid support is a bead. Beads may include one or more types of solid, porous, or hollow spheres, balls, bearings, cylinders, or other similar structures to which nucleic acids can be immobilized (e.g., covalently or non-covalently). Beads may be composed of, for example, plastic, ceramic, metal, polymeric materials, or any combination thereof. Beads may be or include discrete particles that are spherical (e.g., microspheres) or have non-spherical or irregular shapes, such as cubes, rectangular prisms, pyramidal, cylindrical, conical, ellipsoidal, or discoidal shapes. In some embodiments, beads may be non-spherical in shape.

[0118] The beads may comprise a variety of materials, including, but not limited to, paramagnetic materials (e.g., magnesium, molybdenum, lithium, and tantalum), superparamagnetic materials (e.g., ferrite (Fe3O4; magnetite) nanoparticles), ferromagnetic materials (e.g., iron, nickel, cobalt, some alloys thereof, and some rare earth metal compounds), ceramic, plastic, glass, polystyrene, silica, methylstyrene, acrylic polymers, titanium, latex, sepharose, agarose, hydrogels, polymers, cellulose, nylon, or any combination thereof.

[0119] In some embodiments, the beads (e.g., beads having labels attached thereto) are hydrogel beads. In some embodiments, the beads comprise a hydrogel.

[0120] Some embodiments disclosed herein include one or more particles (e.g., beads). Each of the particles may include a plurality of oligonucleotides (e.g., barcodes). Each of the plurality of oligonucleotides may include a barcode sequence (e.g., a molecular tag sequence), a cell tag, and a target binding region (e.g., an oligo(dT) sequence, a gene-specific sequence, a random multimer, or a combination thereof). The cell tag sequence of each of the plurality of oligonucleotides may be the same. The cell tag sequences of oligonucleotides on different particles may be different so that the oligonucleotides on different particles can be identified. The number of different cell tag sequences may vary in different implementations. In some embodiments, the number of cell labeling sequences is 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 , 10 7 , 10 8 , 10 9 , a number or range between, or exceeding, any two of these values, or an approximation of such a value. In some embodiments, the number of cell labeling sequences is at least or at most 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 , 10 7 , 10 8 or 10 9In some embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 or more of the plurality of particles comprise oligonucleotides having the same cellular sequence. In some embodiments, at most 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10% or more of the plurality of particles comprise oligonucleotides having the same cellular sequence. In some embodiments, none of the plurality of particles comprise the same cellular targeting sequence.

[0121] The multiple oligonucleotides on each particle can include different barcode sequences (e.g., molecular labels). In some embodiments, the number of barcode sequences is 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 , 10 7 , 10 8 , 10 9 In some embodiments, the number of barcode sequences is at least or at most 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 , 10 7 , 10 8 or 10 9For example, at least 100 of the plurality of oligonucleotides may comprise different barcode sequences. As another example, in a single particle, at least 100, 500, 1000, 5000, 10000, 15000, 20000, 50000, or a number or range between any two of these values, or more, of the plurality of oligonucleotides may comprise different barcode sequences. Some embodiments provide a plurality of particles comprising barcodes. In some embodiments, the ratio of the presence (or copies or number) of labeled targets to distinct barcode sequences can be at least 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 1:30, 1:40, 1:50, 1:60, 1:70, 1:80, 1:90, or more. In some embodiments, each of the plurality of oligonucleotides further comprises a sample label, a universal label, or both. The particles can be, for example, nanoparticles or microparticles.

[0122] The size of the beads can vary. For example, the diameter of the beads can range from 0.1 micrometers to 50 micrometers. In some embodiments, the diameter of the beads can be 0.1, 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 micrometers, or a number or range between any two of these values, or an approximation of such value.

[0123] The diameter of a bead can be related to the diameter of a well in a substrate. In some embodiments, the diameter of a bead can be 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100% longer or shorter than the diameter of the well, or a number or range between any two of these values, or an approximation of such a value. The diameter of a bead can be related to the diameter of a cell (e.g., a single cell confined in a well in a substrate). In some embodiments, the diameter of a bead can be at least or at most 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% longer or shorter than the diameter of the well. The diameter of a bead can be related to the diameter of a cell (e.g., a single cell confined in a well in a substrate). In some embodiments, the diameter of a bead can be 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 250%, 300% longer or shorter than the diameter of a cell, or a number or range between any two of these values ​​or an approximation of such a value. In some embodiments, the diameter of a bead can be at least or at most 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 250%, or 300% longer or shorter than the diameter of a cell.

[0124] The beads can be bound and / or embedded in a substrate. The beads can be bound and / or embedded in a gel, hydrogel, polymer, and / or matrix. The spatial location of the beads within the substrate (e.g., gel, matrix, scaffold, or polymer) can be identified using spatial labels present in barcodes on the beads, which can serve as location addresses.

[0125] Examples of beads may include, but are not limited to, streptavidin beads, agarose beads, magnetic beads, Dynabeads®, MACS® microbeads, antibody-conjugated beads (e.g., anti-immunoglobulin microbeads), protein A-conjugated beads, protein G-conjugated beads, protein A / G-conjugated beads, protein L-conjugated beads, oligo(dT)-conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, and BcMag™ carboxyl-terminated magnetic beads.

[0126] The beads may be associated with (e.g., impregnated with) quantum dots or fluorescent dyes to fluoresce in one fluorescent optical channel or multiple optical channels. The beads may be associated with iron oxide or chromium oxide to make them paramagnetic or ferromagnetic. The beads may be identifiable. For example, the beads may be imaged using a camera. The beads may have a detectable code associated with them. For example, the beads may include a barcode. The beads may change size, for example, due to swelling in an organic or inorganic solution. The beads may be hydrophobic. The beads may be hydrophilic. The beads may be biocompatible.

[0127] The solid support (e.g., a bead) can be visualized. The solid support can include a visualization tag (e.g., a fluorescent dye). The solid support (e.g., a bead) can be etched with an identifier (e.g., a number). The identifier can be visualized by imaging the bead.

[0128] A solid support can comprise an insoluble, semi-soluble, or insoluble material. A solid support can be referred to as "functionalized" if it contains a linker, backbone, building block, or other reactive moiety attached thereto, while a solid support can be "non-functionalized" if it does not contain such a reactive moiety attached thereto. A solid support can be used free in solution, such as in a microtiter well format; in a flow-through format, such as in a column; or in a dipstick format.

[0129] The solid support may comprise a membrane, paper, plastic, coated surface, flat surface, glass, slide, chip, or any combination thereof. The solid support may take the form of a resin, gel, microsphere, or other geometric shape. The solid support may comprise a silica chip, microparticle, nanoparticle, plate, array, capillary tube, flat support such as a glass fiber filter, glass surface, metal surface (steel, gold, silver, aluminum, silicon, and copper), glass support, plastic support, silicon support, chip, filter, membrane, microwell plate, slide, plastic material (including multiwell plates or membranes (e.g., formed of polyethylene, polypropylene, polyamide, polyvinylidene difluoride)), and / or a wafer, comb, pin, or needle (e.g., an array of pins suitable for combinatorial synthesis or analysis), or beads in a series of depressions or nanoliter wells on a flat surface such as a wafer (e.g., a silicon wafer), a wafer with depressions with or without a filter bottom.

[0130] The solid support may comprise a polymer matrix (e.g., a gel, a hydrogel). The polymer matrix may be capable of penetrating intracellular spaces (e.g., around organelles). The polymer matrix may also be capable of being transported throughout the circulatory system.

[0131] Substrates and microwell arrays As used herein, a substrate may refer to a type of solid support. A substrate may refer to a solid support that may include a barcode or stochastic barcode of the present disclosure. A substrate may include, for example, a plurality of microwells. For example, a substrate may be a well array including two or more microwells. In some embodiments, a microwell may include a small reaction chamber of a defined volume. In some embodiments, a microwell may confine one or more cells. In some embodiments, a microwell may confine only one cell. In some embodiments, a microwell may confine one or more solid supports. In some embodiments, a microwell may confine only one solid support. In some embodiments, a microwell confines a single cell and a single solid support (e.g., a bead). A microwell may include a barcode reagent of the present disclosure.

[0132] Barcoding methods The present disclosure provides methods for estimating the number of unique targets at unique locations in a bodily sample (e.g., tissue, organ, tumor, cell). The methods may include placing a barcode (e.g., a probabilistic barcode) in proximity to the sample, lysing the sample, associating the unique targets with the barcode, amplifying the targets, and / or digitally counting the targets. The methods may further include analyzing and / or visualizing information obtained from the spatial labeling of the barcode. In some embodiments, the methods include visualizing multiple targets in the sample. Mapping the multiple targets to a map of the sample may include creating a two-dimensional or three-dimensional map of the sample. The two-dimensional and three-dimensional maps may be created before or after barcoding (e.g., probabilistic barcoding) the multiple targets in the sample. Visualizing multiple targets in the sample may include mapping the multiple targets to a map of the sample. Mapping the multiple targets to a map of the sample may include creating a two-dimensional or three-dimensional map of the sample. The two-dimensional and three-dimensional maps can be generated before or after barcoding multiple targets in a sample. In some embodiments, the two-dimensional and three-dimensional maps can be generated before or after lysing the sample. Lysing the sample before or after generating the two-dimensional or three-dimensional map can include heating the sample, contacting the sample with a detergent, changing the pH of the sample, or any combination thereof.

[0133] In some embodiments, barcoding the plurality of targets comprises hybridizing the plurality of barcodes to the plurality of targets to generate barcoded targets (e.g., stochastically barcoded targets). Barcoding the plurality of targets may comprise creating an indexed library of barcoded targets. Creating an indexed library of barcoded targets may be performed using a solid support comprising a plurality of barcodes (e.g., stochastic barcodes).

[0134] Contacting the sample and barcode The present disclosure provides methods for contacting a sample (e.g., cells) with a substrate of the present disclosure. For example, a sample including cells, an organ, or a tissue slice can be contacted with a barcode (e.g., a stochastic barcode). The cells can be contacted, for example, by gravity flow, in which case the cells can settle to form a monolayer. The sample can be a tissue slice. The slice can be disposed on a substrate. The sample can be one-dimensional (e.g., forming a planar surface). The sample (e.g., cells) can be spread across the substrate, for example, by growing / culturing the cells on the substrate.

[0135] When the barcode is in proximity to the target, the target can hybridize to the barcode. The barcodes can be contacted in a non-depleting ratio so that each unique target can bind to a unique barcode of the present disclosure. To ensure efficient binding between the target and the barcode, the target can be cross-linked to the barcode.

[0136] Cell lysis After partitioning the cells and barcodes, the cells can be lysed to release the target molecule. Cell lysis can be achieved by any of a variety of means, such as chemical or biochemical means, osmotic shock, or thermal lysis, mechanical lysis, or optical lysis. Cells can be lysed by adding a cell lysis buffer containing a detergent (e.g., SDS, Li dodecyl sulfate, Triton X-100, Tween-20, or NP-40), an organic solvent (e.g., methanol or acetone), or a digestive enzyme (e.g., proteinase K, pepsin, or trypsin), or any combination thereof. To improve target and barcode association, the diffusion rate of the target molecule can be altered, for example, by lowering the temperature and / or increasing the viscosity of the lysate.

[0137] In some embodiments, the sample can be lysed using filter paper, which can be soaked with a lysis buffer over the filter paper, and the filter paper can be applied to the sample with pressure, which can promote lysis of the sample and hybridization of the sample's targets to the substrate.

[0138] In some embodiments, lysis may be performed by mechanical, thermal, optical, and / or chemical lysis. Chemical lysis may include the use of digestive enzymes such as proteinase K, pepsin, and trypsin. Lysis may be performed by adding a lysis buffer to the substrate. The lysis buffer may include Tris-HCl. The lysis buffer may include at least about 0.01, 0.05, 0.1, 0.5, or 1 M or more Tris-HCl. The lysis buffer may include at most about 0.01, 0.05, 0.1, 0.5, or 1 M or more Tris-HCl. The lysis buffer may include about 0.1 M Tris-HCl. The pH of the lysis buffer may be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. The pH of the lysis buffer may be at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. In some embodiments, the pH of the lysis buffer is about 7.5. The lysis buffer may include a salt (e.g., LiCl). The concentration of the salt in the lysis buffer may be at least about 0.1, 0.5, or 1 M or more. The concentration of the salt in the lysis buffer may be at most about 0.1, 0.5, or 1 M or more. In some embodiments, the concentration of the salt in the lysis buffer is about 0.5 M. The lysis buffer may include a detergent (e.g., SDS, Li dodecyl sulfate, triton X, tween, NP-40). The concentration of the detergent in the lysis buffer may be at least about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, or 7% or more. The concentration of the detergent in the lysis buffer can be at most about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, or 7% or more. In some embodiments, the concentration of the detergent in the lysis buffer is about 1% Li dodecyl sulfate. The time used in the method for lysis can depend on the amount of detergent used. In some embodiments, the more detergent used, the shorter the time required for lysis. The lysis buffer can include a chelating agent (e.g., EDTA, EGTA).The concentration of the chelating agent in the lysis buffer may be at least about 1, 5, 10, 15, 20, 25, or 30 mM or more. The concentration of the chelating agent in the lysis buffer may be at most about 1, 5, 10, 15, 20, 25, or 30 mM or more. In some embodiments, the concentration of the chelating agent in the lysis buffer is about 10 mM. The lysis buffer may include a reducing reagent (e.g., β-mercaptoethanol, DTT). The concentration of the reducing reagent in the lysis buffer may be at least about 1, 5, 10, 15, or 20 mM or more. The concentration of the reducing reagent in the lysis buffer may be at most about 1, 5, 10, 15, or 20 mM or more. In some embodiments, the concentration of the reducing reagent in the lysis buffer is about 5 mM. In some embodiments, the lysis buffer may comprise about 0.1 M Tris-HCl, about pH 7.5, about 0.5 M LiCl, about 1% lithium dodecyl sulfate, about 10 mM EDTA, and about 5 mM DTT.

[0139] Lysing may be performed at a temperature of about 4, 10, 15, 20, 25, or 30° C. Lysing may be performed for about 1, 5, 10, 15, or 20 minutes or more. Lysed cells may contain at least about 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000 or more target nucleic acid molecules. Lysed cells may contain at most about 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000 or more target nucleic acid molecules.

[0140] Binding of barcodes to target nucleic acid molecules After cell lysis and release of nucleic acid molecules therefrom, the nucleic acid molecules can be randomly associated with the barcodes on the co-localized solid support. Association can involve hybridization of the target recognition region of the barcode to a complementary portion of the target nucleic acid molecule (e.g., the oligo(dT) of the barcode can interact with the poly(A) tail of the target). Assay conditions (e.g., buffer pH, ionic strength, temperature, etc.) used for hybridization can be selected to promote the formation of specific, stable hybrids. In some embodiments, nucleic acid molecules released from lysed cells can be associated with (e.g., hybridized to) multiple probes on a substrate. If the probes contain oligo(dT), mRNA molecules can hybridize to the probes and be reverse transcribed. The oligo(dT) portion of the oligonucleotide can act as a primer for first-strand synthesis of cDNA molecules. For example, in the non-limiting example of barcoding shown in Figure 2, block 216, mRNA molecules can hybridize to barcodes on beads. For example, a single-stranded nucleotide fragment can hybridize to the target binding region of the barcode.

[0141] The binding may further include ligating the target recognition region of the barcode and a portion of the target nucleic acid molecule. For example, the target binding region may include a nucleic acid sequence capable of specific hybridization to a restriction site overhang (e.g., an EcoRI sticky end overhang). The assay procedure may further include treating the target nucleic acid with a restriction enzyme (e.g., EcoRI) to generate the restriction site overhang. The barcode may then be ligated to any nucleic acid molecule containing a sequence complementary to the restriction site overhang. A ligase (e.g., T4 DNA ligase) may be used to link the two fragments.

[0142] For example, in a non-limiting example of barcoding shown in Figure 2, block 220, labeled targets (e.g., target-barcode molecules) from multiple cells (or multiple samples) can then be pooled, e.g., in a tube. The labeled targets can be pooled, e.g., by collecting beads to which the barcodes and / or target-barcode molecules are bound.

[0143] Solid support-based collection recovery of bound target-barcode molecules can be achieved through the use of magnetic beads and an externally applied magnetic field. After the target-barcode molecules are pooled, all further processing can proceed in a single reaction vessel. Further processing can include, for example, reverse transcription reactions, amplification reactions, cleavage reactions, dissociation reactions, and / or nucleic acid extension reactions. Further processing reactions can be performed within microwells, i.e., without first pooling the labeled target nucleic acid molecules from multiple cells.

[0144] Reverse transcription The present disclosure provides a method for generating a target-barcode conjugate using reverse transcription (e.g., at block 224 of Figure 2). The target-barcode conjugate can include a barcode and a complementary sequence of all or part of a target nucleic acid (i.e., a barcoded cDNA molecule, such as a stochastically barcoded cDNA molecule). Reverse transcription of the associated RNA molecule can occur by adding a reverse transcription primer along with a reverse transcriptase. The reverse transcription primer can be an oligo(dT) primer, a random hexanucleotide primer, or a target-specific oligonucleotide primer. The oligo(dT) primer can be 12-18 nucleotides in length or approximately such nucleotides and binds to the endogenous poly(A) tail at the 3' end of mammalian mRNA. The random hexanucleotide primer can bind to the mRNA at various complementary sites. The target-specific oligonucleotide primer typically selectively primes the mRNA of interest.

[0145] In some embodiments, reverse transcription of the labeled RNA molecule can occur by adding a reverse transcription primer. In some embodiments, the reverse transcription primer is an oligo(dT) primer, a random hexanucleotide primer, or a target-specific oligonucleotide primer. Generally, oligo(dT) primers are 12-18 nucleotides in length and bind to the endogenous poly(A) tail at the 3' end of mammalian mRNAs. Random hexanucleotide primers can bind to mRNAs at various complementary sites. Target-specific oligonucleotide primers typically selectively prime the mRNA of interest.

[0146] Reverse transcription can be performed repeatedly to generate multiple labeled cDNA molecules. The methods disclosed herein can include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 reverse transcription reactions. The methods can include performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 reverse transcription reactions.

[0147] amplification One or more nucleic acid amplification reactions (e.g., at block 228 of FIG. 2 ) can be performed to generate multiple copies of the labeled target nucleic acid molecule. Amplification can be performed in a multiplexed manner, where multiple target nucleic acid sequences are amplified simultaneously. The amplification reaction can be used to add sequencing adapters to the nucleic acid molecule. The amplification reaction can include amplifying at least a portion of the sample label, if present. The amplification reaction can include amplifying at least a portion of the cell label and / or barcode sequence (e.g., molecular label). The amplification reaction can include amplifying at least a portion of the sample tag, cell label, spatial label, barcode sequence (e.g., molecular label), target nucleic acid, or a combination thereof. The amplification reaction may include amplifying 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 100%, or a range or number between any two of these values ​​of the plurality of nucleic acids. The method may further include performing one or more cDNA synthesis reactions to generate one or more cDNA copies of the target-barcode molecule comprising the sample label, cell label, spatial label, and / or barcode sequence (e.g., molecular label).

[0148] In some embodiments, amplification can be performed using polymerase chain reaction (PCR). As used herein, PCR can refer to a reaction for in vitro amplification of specific DNA sequences by simultaneous primer extension of complementary strands of DNA. As used herein, PCR can encompass derivatives of the reaction, including, but not limited to, RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, and assembly PCR.

[0149] Amplification of the labeled nucleic acid may include non-PCR-based methods. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or circle-circle amplification. Other non-PCR-based amplification methods include DNA-dependent RNA polymerase-driven RNA transcription amplification or multiple cycles of RNA-directed DNA synthesis and transcription to amplify DNA or RNA targets, ligase chain reaction (LCR) and Qβ replicase (Qβ) methods, the use of palindromic probes, strand displacement amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods in which a primer is hybridized to a nucleic acid sequence and the resulting duplex is cleaved before extension and amplification, strand displacement amplification using a nucleic acid polymerase lacking 5' exonuclease activity, rolling circle amplification, and branched extension amplification (RAM). In some embodiments, amplification does not produce circularized transcripts.

[0150] In some embodiments, the methods disclosed herein further include performing a polymerase chain reaction on the labeled nucleic acid (e.g., labeled RNA, labeled DNA, labeled cDNA) to generate a labeled amplicon (e.g., a stochastically labeled amplicon). The labeled amplicon can be a double-stranded molecule. The double-stranded molecule can include a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized to a DNA molecule. One or both strands of the double-stranded molecule can include a sample label, a spatial label, a cell label, and / or a barcode sequence (e.g., a molecular label). The labeled amplicon can be a single-stranded molecule. The single-stranded molecule can include DNA, RNA, or a combination thereof. The nucleic acids of the present disclosure can include synthetic or modified nucleic acids.

[0151] Amplification may include the use of one or more non-natural nucleotides. Non-natural nucleotides may include photolabile or trigger nucleotides. Examples of non-natural nucleotides may include, but are not limited to, peptide nucleic acids (PNAs), morpholinos, locked nucleic acids (LNAs), glycol nucleic acids (GNAs), and threose nucleic acids (TNAs). Non-natural nucleotides may be added to one or more cycles of the amplification reaction. The addition of non-natural nucleotides may be used to identify products at specific cycles or time points of the amplification reaction.

[0152] The step of performing one or more amplification reactions may include the use of one or more primers. The one or more primers may comprise, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides or more. The one or more primers may comprise at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides or more. The one or more primers may comprise 12 to fewer than 15 nucleotides. The one or more primers may anneal to at least a portion of the multiple labeled targets (e.g., stochastically labeled targets). The one or more primers may anneal to the 3' or 5' ends of the multiple labeled targets. The one or more primers may anneal to an internal region of the multiple labeled targets. The internal region can be at least about 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of the multiple labeled targets. The one or more primers can comprise a fixed panel of primers. The one or more primers can include at least one or more custom primers. The one or more primers may include at least one or more control primers. The one or more primers may include at least one or more gene-specific primers.

[0153] The one or more primers may include a universal primer. The universal primer may anneal to a universal primer binding site. The one or more custom primers may anneal to a first sample label, a second sample label, a spatial label, a cell label, a barcode sequence (e.g., a molecular label), a target, or any combination thereof. The one or more primers may include a universal primer and a custom primer. The custom primers may be designed to amplify one or more targets. The targets may comprise a subset of all nucleic acids in one or more samples. The targets may comprise a subset of all labeled targets in one or more samples. The one or more primers may include at least 96 or more custom primers. The one or more primers may include at least 960 or more custom primers. The one or more primers may include at least 9600 or more custom primers. The one or more custom primers may anneal to two or more different labeled nucleic acids. The two or more different labeled nucleic acids may correspond to one or more genes.

[0154] Any amplification scheme can be used in the disclosed methods. For example, in one scheme, a first round of PCR can amplify molecules bound to beads using a gene-specific primer and a primer for the universal Illumina sequencing primer 1 sequence. A second round of PCR can amplify the first PCR product using a nested gene-specific primer and a primer for the universal Illumina sequencing primer 1 sequence flanked by Illumina sequencing primer 2 sequences. A third round of PCR adds P5, P7, and a sample index to convert the PCR products into an Illumina sequencing library. Sequencing using 150 bp x 2 sequencing can reveal cell label and barcode sequences (e.g., molecular labels) on read 1, genes on read 2, and a sample index on read index 1.

[0155] In some embodiments, nucleic acids can be removed from a substrate using chemical cleavage. For example, chemical groups or modified bases present in the nucleic acid can be used to facilitate its removal from a solid support. For example, enzymes can be used to remove nucleic acids from a substrate. For example, nucleic acids can be removed from a substrate by restriction endonuclease (also referred to herein as "restriction enzyme") digestion. For example, treatment of nucleic acids containing dUTP or ddUTP with uracil-d-glycosylase (UDG) can be used to remove nucleic acids from a substrate. For example, nucleic acids can be removed from a substrate using an enzyme that performs nucleotide excision, such as a base excision repair enzyme, e.g., an apurinic / apyrimidinic (AP) endonuclease. In some embodiments, nucleic acids can be removed from a substrate using a photocleavable group and light. In some embodiments, a cleavable linker can be used to remove nucleic acids from a substrate. For example, the cleavable linker can comprise at least one of biotin / avidin, biotin / streptavidin, biotin / neutravidin, Ig-Protein A, a photolabile linker, an acid or base labile linker group, or an aptamer.

[0156] If the probe is gene-specific, the molecule can be hybridized to the probe and reverse transcribed and / or amplified. In some embodiments, after the nucleic acid is synthesized (e.g., reverse transcribed), it can be amplified. Amplification can be performed in a multiplexed manner, where multiple target nucleic acid sequences are amplified simultaneously. Amplification can add sequencing adapters to the nucleic acid.

[0157] In some embodiments, amplification can be performed on the substrate, for example, using bridge amplification. The cDNA can be homopolymer tailed to generate ends compatible with bridge amplification on the substrate using an oligo(dT) probe. In bridge amplification, a primer complementary to the 3' end of the template nucleic acid can be the first primer of each pair covalently attached to a solid particle. When a sample containing the template nucleic acid is contacted with the particle and one thermal cycle is performed, the template molecule can anneal to the first primer, and the first primer can be extended in the forward direction by the addition of nucleotides to form a double-stranded molecule consisting of the template molecule and a newly formed DNA strand complementary to the template. In the heating step of the next cycle, the double-stranded molecule can be denatured, releasing the template molecule from the particle and leaving a complementary DNA strand attached to the particle via the first primer. In the annealing stage of the subsequent annealing and extension step, the complementary strand can hybridize to a second primer complementary to the segment of the complementary strand at the position removed from the first primer. This hybridization allows the complementary strand to form a bridge between the first and second primers, covalently bound to the first primer and hybridized to the second primer. In the extension step, the second primer can be extended in the opposite direction by adding nucleotides to the same reaction mixture, thereby converting the bridge into a double-stranded bridge. The next cycle then begins, and the double-stranded bridge is denatured to produce two single-stranded nucleic acid molecules, each with one end bound to the particle surface via the first and second primers and the other end unbound, respectively. In the annealing and extension step of this second cycle, each strand can hybridize to additional, previously unused, complementary primers on the same particle to form a new single-stranded bridge. The two previously unused hybridized primers can then be extended to convert the two new bridges into double-stranded bridges.

[0158] The amplification reaction can include amplifying at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97% or 100% of the plurality of nucleic acids.

[0159] Amplification of the labeled nucleic acid may include PCR-based or non-PCR-based methods. Amplification of the labeled nucleic acid may include exponential amplification of the labeled nucleic acid. Amplification of the labeled nucleic acid may include linear amplification of the labeled nucleic acid. Amplification may be performed by polymerase chain reaction (PCR). PCR may refer to a reaction for in vitro amplification of specific DNA sequences by simultaneous primer extension of complementary strands of DNA. PCR may encompass derivatives of the reaction, including, but not limited to, RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, suppression PCR, semi-suppressive PCR, and assembly PCR.

[0160] In some embodiments, amplification of the labeled nucleic acid comprises a non-PCR-based method. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or circle-circle amplification. Other non-PCR-based amplification methods include DNA-dependent RNA polymerase-driven RNA transcription amplification or multiple cycles of RNA-directed DNA synthesis and transcription to amplify DNA or RNA targets, ligase chain reaction (LCR), Qβ replicase (Qβ), the use of palindromic probes, strand displacement amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods in which a primer is hybridized to a nucleic acid sequence and the resulting duplex is cleaved before extension and amplification, strand displacement amplification using a nucleic acid polymerase lacking 5' exonuclease activity, rolling circle amplification, and / or branched extension amplification (RAM).

[0161] In some embodiments, the methods disclosed herein further comprise performing a nested polymerase chain reaction on the amplified amplicon (e.g., target). The amplicon may be a double-stranded molecule. The double-stranded molecule may comprise a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized to a DNA molecule. One or both strands of the double-stranded molecule may comprise a sample tag or molecular identifier label. Alternatively, the amplicon may be a single-stranded molecule. The single-stranded molecule may comprise DNA, RNA, or a combination thereof. The nucleic acids described herein may include synthetic or modified nucleic acids.

[0162] In some embodiments, the methods include repeatedly amplifying a labeled nucleic acid to generate multiple amplicons. The methods disclosed herein may include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amplification reactions. Alternatively, the methods include performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 amplification reactions.

[0163] The amplification may further include adding one or more control nucleic acids to one or more samples containing the plurality of nucleic acids. The amplification may further include adding one or more control nucleic acids to the plurality of nucleic acids. The control nucleic acids may include a control label.

[0164] Amplification may involve the use of one or more non-natural nucleotides. Non-natural nucleotides may include photolabile and / or trigger nucleotides. Examples of non-natural nucleotides include, but are not limited to, peptide nucleic acids (PNAs), morpholinos, locked nucleic acids (LNAs), glycol nucleic acids (GNAs), and threose nucleic acids (TNAs). Non-natural nucleotides may be added in one or more cycles of the amplification reaction. The addition of non-natural nucleotides may be used to identify products at specific cycles or time points of the amplification reaction.

[0165] The step of performing one or more amplification reactions may include the use of one or more primers. The one or more primers may comprise one or more oligonucleotides. The one or more oligonucleotides may comprise at least about 7 to 9 nucleotides. The one or more oligonucleotides may comprise less than 12 to 15 nucleotides. The one or more primers may anneal to at least a portion of the plurality of labeled nucleic acids. The one or more primers may anneal to the 3' and / or 5' ends of the plurality of labeled nucleic acids. The one or more primers may anneal to an internal region of the plurality of labeled nucleic acids. The internal region can be at least about 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of the plurality of labeled nucleic acids. The one or more primers can comprise a fixed panel of primers. The one or more primers can comprise at least one or more custom primers. The one or more primers may include at least one or more control primers. The one or more primers may include at least one or more housekeeping gene primers. The one or more primers may include a universal primer. The universal primer may anneal to a universal primer binding site. The one or more custom primers may anneal to a first sample tag, a second sample tag, a molecular identifier label, a nucleic acid, or a product thereof. The one or more primers may include a universal primer and a custom primer. The custom primer may be designed to amplify one or more target nucleic acids. The target nucleic acids may comprise a subset of the total nucleic acids in one or more samples. In some embodiments, the primers are probes attached to the array of the present disclosure.

[0166] In some embodiments, barcoding (e.g., stochastically barcoding) a plurality of targets in a sample further comprises generating an indexed library of barcoded targets (e.g., stochastically barcoded targets) or barcoded fragments of the targets. The barcode sequences of different barcodes (e.g., molecular labels of different stochastic barcodes) can differ from one another. Generating an indexed library of barcoded targets comprises generating a plurality of indexed polynucleotides from the plurality of targets in the sample. For example, for an indexed library of barcoded targets comprising a first indexed target and a second indexed target, the labeled region of the first indexed polynucleotide can differ from the labeled region of the second indexed polynucleotide by, or by about at least or at most, such a value, or a number or range of nucleotides between any two of these values. In some embodiments, the step of generating an indexed library of barcoded targets includes contacting a plurality of targets, e.g., mRNA molecules, with a plurality of oligonucleotides each comprising a poly(T) region and a label region; and performing first-strand synthesis using a reverse transcriptase to generate single-stranded, labeled cDNA molecules, each comprising a cDNA region and a label region, wherein the plurality of targets comprises at least two mRNA molecules of different sequences and the plurality of oligonucleotides comprises at least two oligonucleotides of different sequences. The step of generating an indexed library of barcoded targets may further include amplifying the single-stranded, labeled cDNA molecules to generate double-stranded, labeled cDNA molecules; and performing nested PCR on the double-stranded, labeled cDNA molecules to generate labeled amplicons. In some embodiments, the method may include generating adapter-labeled amplicons.

[0167] Barcoding (e.g., probabilistic barcoding) can involve using nucleic acid barcodes or tags to label individual nucleic acid (e.g., DNA or RNA) molecules. In some embodiments, it involves adding DNA barcodes or tags to cDNA molecules as they are generated from mRNA. Nested PCR can be performed to minimize PCR amplification bias. Adapters can be added for sequencing, e.g., using next-generation sequencing (NGS). Sequencing results can be used to determine the sequence of cellular labels, molecular labels, and nucleotide fragments of one or more copies of the target, e.g., in block 232 of FIG. 2.

[0168] 3 is a schematic diagram illustrating a non-limiting, exemplary process for generating an indexed library of barcoded targets (e.g., stochastically barcoded targets), such as barcoded mRNAs or fragments thereof. As shown in step 1, the reverse transcription process can encode each mRNA molecule containing a unique molecular tag sequence, a cellular tag sequence, and a universal PCR site. In particular, an RNA molecule 302 can be reverse transcribed by hybridization (e.g., stochastic hybridization) of a set of barcodes (e.g., stochastic barcodes) 310 to a poly(A) tail region 308 of the RNA molecule 302 to generate labeled cDNA molecules 304 containing cDNA regions 306. Each of the barcodes 310 can include a target binding region, e.g., a poly(dT) region 312, a tag region 314 (e.g., a barcode sequence or molecule), and a universal PCR region 316.

[0169] In some embodiments, the cell label sequence can comprise 3 to 20 nucleotides. In some embodiments, the molecular label sequence can comprise 3 to 20 nucleotides. In some embodiments, each of the plurality of stochastic barcodes further comprises one or more of a universal label and a cell label, wherein the universal label is the same for the plurality of stochastic barcodes on the solid support and the cell label is the same for the plurality of stochastic barcodes on the solid support. In some embodiments, the universal label can comprise 3 to 20 nucleotides. In some embodiments, the cell label comprises 3 to 20 nucleotides.

[0170] In some embodiments, label region 314 can include a barcode sequence or molecular label 318 and a cell label 320. In some embodiments, label region 314 can include one or more of a universal label, a dimensional label, and a cell label. Barcode sequence or molecular label 318 can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, or approximately at least or at most such a number of nucleotides in length, or a number or range of nucleotides in length between any of these values. Cell label 320 can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, or approximately at least or at most such a number of nucleotides in length, or a number or range of nucleotides in length between any of these values. The universal label can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, or about at least or at most such values, or a number or range of nucleotides in length between any of these values. The universal label can be the same for multiple probabilistic barcodes on a solid support, and the cell label is the same for multiple probabilistic barcodes on a solid support. The dimensional labels can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, or about at least or at most such values, or a number or range of nucleotides in length between any of these values.

[0171] In some embodiments, label region 314 can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 different labels, or about such a number of different labels, or at least such a number of different labels, or a number or range of different labels between any of these values, such as barcode sequences or molecular labels 318 and cellular labels 320. Each label can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, or about at least such a number of different labels, or a number or range of different labels between any of these values. The set of barcodes or probabilistic barcodes 310 may be: 10, 20, 40, 50, 70, 80, 90, 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 The set of barcodes or probabilistic barcodes 310 may include at least about or at most such number of barcodes or probabilistic barcodes 310, or a number or range of barcodes or probabilistic barcodes 310 between any of these values. The set of barcodes or probabilistic barcodes 310 may, for example, each include a unique labeled region 314. The labeled cDNA molecules 304 may be purified to remove excess barcodes or probabilistic barcodes 310. Purification may include Ampure bead purification.

[0172] As shown in step 2, the products from the reverse transcription process in step 1 can be pooled in one tube and PCR amplified using a first pool of PCR primers and a first universal PCR primer. Pooling is possible due to the uniquely labeled region 314. In particular, the labeled cDNA molecules 304 can be amplified to generate nested PCR-labeled amplicons 322. The amplification can include multiplex PCR amplification. The amplification can include multiplex PCR amplification using 96 multiplex primers in a single reaction volume. In some embodiments, the multiplex PCR amplification can be performed using 10, 20, 40, 50, 70, 80, 90, 10, 25, 30, 45, 50, 60, 75, 80, 90, 100, 150, 250, 300, 450, 500, 600, 750, 800, 900, 1500, 1500, 2500, 3000, 4500, 5000, 6000, 15000, 25000, 30000, 45000, 50000, 50000, 60000, 75000, 8000, 9000, 15000, 15000, 15000, 25000, 30000, 45000, 50 ...0, 250000, 30000, 45000, 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 The amplification may include using a first PCR primer pool 324 that includes custom primers 326A-C that target specific genes and a universal primer 328. The custom primer 326 may hybridize to a region within the cDNA portion 306' of the labeled cDNA molecule 304. The universal primer 328 may hybridize to the universal PCR region 316 of the labeled cDNA molecule 304.

[0173] As shown in step 3 of Figure 3, the product from the PCR amplification in step 2 can be amplified using a nested PCR primer pool and a second universal PCR primer. Nested PCR can minimize PCR amplification bias. In particular, nested PCR-labeled amplicons 322 can be further amplified by nested PCR. Nested PCR can include multiplex PCR including a nested PCR primer pool 330 of nested PCR primers 332a-c and a second universal PCR primer 328' in a single reaction volume. The nested PCR primer pool 328 may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 different nested PCR primers 330, or may include approximately such a number of different nested PCR primers 330, or may include at least or at most such a number of different nested PCR primers 330, or a number or range between any of these values. The nested PCR primers 332 may contain an adaptor 334 and hybridize to a region within the cDNA portion 306" of the labeled amplicon 322. The universal primer 328' may contain an adaptor 336 and hybridize to the universal PCR region 316 of the labeled amplicon 322. Thus, step 3 generates adapter-labeled amplicon 338. In some embodiments, nested PCR primer 332 and second universal PCR primer 328' may not contain adapters 334 and 336. Instead, adapters 334 and 336 may be ligated to the product of the nested PCR to generate adapter-labeled amplicon 338.

[0174] As shown in step 4, the PCR products from step 3 can be PCR amplified for sequencing using library amplification primers. In particular, adapters 334 and 336 can be used to perform one or more additional assays on adapter-labeled amplicons 338. Adapters 334 and 336 can be hybridized to primers 340 and 342. One or more of primers 340 and 342 can be PCR amplification primers. One or more of primers 340 and 342 can be sequencing primers. One or more of adapters 334 and 336 can be used for further amplification of adapter-labeled amplicons 338. One or more of adapters 334 and 336 can be used for sequencing of adapter-labeled amplicons 338. Primer 342 can contain plate index 344 so that amplicons or probability barcodes 310 generated using the same set of barcodes can be sequenced in a single sequencing reaction using next-generation sequencing (NGS).

[0175] Multi-omics analysis Disclosed herein are embodiments of methods for high-throughput sample analysis. The methods can be used with any sample analysis platform or system for partitioning single cells using single particles, such as droplets (e.g., Chromium™ Single Cell 3' Solution (10X Genomics, San Francisco, CA)), microwells (e.g., Rhapsody™ Assay (Becton, Dickinson and Company, Franklin Lakes, NJ)), microfluidic chambers, and patterned substrate-based platforms and systems. The methods can capture multi-omic information, including genome, genome accessibility (e.g., chromatin accessibility), and methylome. The methods can be used with methods for transcriptomics analysis, proteomics analysis, and / or sample tracking. The use of barcoding for proteomic analysis is described in U.S. Patent Application No. 15 / 715,028 (published as U.S. Patent Application Publication No. 2018 / 0088112), the entire contents of which are incorporated herein by reference. The use of barcoding for sample tracking is described in U.S. Patent Application No. 15 / 937,713 (published as U.S. Patent Application Publication No. 2018 / 0346970), the entire contents of which are incorporated herein by reference. In some embodiments, multi-omic information, such as single-cell genomics, chromatin accessibility, methylomics, transcriptomics, and proteomics, is obtained using barcoding.

[0176] In some embodiments, the method includes adding sequences complementary to those of capture probes bearing cellular and molecular labels or indicators to the ends of genomic DNA fragments. For example, a poly(dA) tail (or any sequence) can be added to a genomic fragment so that the genomic fragment can be captured by an oligo(dT) probe (or a sequence complementary to the added sequence) flanked by cellular and molecular barcodes. This method can be used to capture all or part of the following from a single cell in a high-throughput manner: genome, methylome, chromatin accessibility, transcriptome, and proteome.

[0177] The method may include sample preparation before loading cellular material into any of these sample analysis systems. For example, using an enzymatic cutter, such as a double-stranded nuclease (such as a transposase, restriction enzyme, or CRISPR-associated protein) described herein, dsDNA (e.g., gDNA) can be fragmented into genomic fragments in fixed cells or nuclei. Restriction enzymes can be used for high-throughput multi-omics sample analysis. For example, the method may include incubating cells with a restriction enzyme (followed by, e.g., removing the restriction enzyme). As another example, the method may include incubating cells with a ligase and an adaptor having a poly(dT) / poly(dA) or poly(dT) / poly(dA) with a T7 promoter sequence flanked by restriction sequences. As yet another example, the capture probe may have the sequence of a restriction site. In this embodiment, the addition of a dT / dA adaptor may not be required. In some embodiments, Cas9 / CRISPR can be used to cleave at specific locations in the genome.

[0178] The cells or nuclei can be fresh or fixed (e.g., cells fixed with a fixative such as an aldehyde, an oxidizing agent, or a Hepes-glutamate buffer-mediated organic solvent protective effect (HOPE) fixative). In some embodiments, the method includes contacting the cells with a nucleic acid reagent described herein. The cells can then be washed to remove excess nucleic acid reagent. As described herein, the nucleic acid reagent can bind to dsDNA in dead cells but not to live cells, such that only dead cells remain labeled with the nucleic acid reagent after washing. In some embodiments, a sequence complementary to a capture probe (e.g., a barcode, such as a stochastic barcode) is then added to each end of the genome fragment. The capture probe can be immobilized on a solid support or in solution. The capture probe in the single-cell transcriptome analysis system can be a poly(dT) sequence. Thus, a poly(dA) sequence can be added to each end of the genome fragment. The cells or nuclei can then be heated or exposed to chemicals to denature the double-stranded genomic fragments with poly(dA) sequences added to each end, and then loaded into a sample analysis system. After cell and / or nuclear lysis, the genomic fragments with the added sequences can be captured by the existing capture probes, similar to how mRNA molecules with poly(A) tails can be captured by the poly(dT) sequences of a capture probe. Reverse transcriptase and / or DNA polymerase can be added to copy (e.g., reverse transcribe) the genomic fragments and attach cellular and molecular labels or indicators to the genomic fragments.

[0179] Disclosed herein are embodiments of methods for sample analysis. Figures 4A-4B show schematic diagrams of non-limiting, exemplary embodiments of a method 400 for high-throughput capture of multi-omic information from single cells. In some embodiments, method 400 includes using a transposome to generate double-stranded DNA fragments having 5' overhangs (or 3' overhangs) that include capture sequences. Method 400 may include contacting 410 double-stranded deoxyribonucleic acid (dsDNA), such as genomic DNA (gDNA), with a transposome 428. Transposome 428 may include a double-stranded nuclease 430 configured to induce double-stranded DNA breaks in the dsDNA-containing structure and two copies 432a, 432b of an adapter with 5' overhangs that include capture sequences (e.g., poly(dT) sequences 434a, 434b). The double-stranded nuclease 430 can be loaded with two copies 432a and 432b of adapters 434a and 434b. Each copy 436a and 436b of the adapter can include a transposon DNA end sequence (e.g., a Tn5 sequence 436a and 436b, or a subsequence thereof). The double-stranded nuclease can be or include a transposase, such as a Tn5 transposase. Contacting 410 the dsDNA (e.g., gDNA) with the transposome 428 can generate multiple overhanging dsDNA fragments 438, each having two copies 432a and 432b of a 5' overhang 434a and 434b.

[0180] In some embodiments, method 400 includes contacting a plurality of overhanging dsDNA fragments (having 5' overhangs) 438 with a polymerase (e.g., at block 412) to generate a plurality of complementary dsDNA fragments, each comprising a sequence 434a', 434b' complementary to at least a portion of 5' overhangs 434a, 434b. Method 400 may also include denaturing (e.g., at block 414) a plurality of complementary dsDNA fragments 440, each comprising a sequence complementary to at least a portion of the 5' overhangs, to generate a plurality of single-stranded DNA (ssDNA) fragments 442, and barcoding (e.g., at block 424) the plurality of ssDNA fragments with a plurality of barcodes 444 to generate a plurality of barcoded ssDNA fragments (e.g., barcoded ssDNA fragments 446 or complementary sequences thereof). At least some (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 10, 100, 1000, 10000, 100000, 1000000, 10000000, or more) of the plurality of barcodes 444 include a cell label 448, a molecular label 450, and a capture sequence 434. The molecular labels 448 of at least two barcodes of the plurality of barcodes 444 may include different molecular label sequences. At least two barcodes of the plurality of barcodes 444 may include a cell label 450 having the same cell label sequence. Method 400 may include obtaining sequencing data of the plurality of barcoded ssDNA fragments 446 (or complementary sequences thereof) and determining information about the dsDNA (e.g., gDNA) based on the sequences (or complementary sequences) of the plurality of ssDNA fragments 446 in the obtained sequencing data.

[0181] Method 400 may include generating DNA fragments from genomic DNA of a cell using a transposome 428 (which may include, for example, a transposase, a restriction endonuclease, and / or a CRISPR-associated protein such as Cas9 or Cas12a). In some embodiments, method 400 may include generating a plurality of nucleic acid fragments from double-stranded deoxyribonucleic acid (dsDNA) of the cell, such as gDNA. For example, the plurality of nucleic acid fragments may not be generated from amplification. As another example, the plurality of nucleic acid fragments may be or include RNA molecules produced by in vitro transcription.

[0182] In some embodiments, each of the plurality of nucleic acid fragments can include a capture sequence 434a, 434b, a complement of the capture sequence, a reverse complement of the capture sequence, or a combination thereof. Method 400 can include barcoding 424 the plurality of nucleic acid fragments using a plurality of barcodes 444 to generate a plurality of barcoded single-stranded deoxyribonucleic acid (ssDNA) fragments 446 (or complementary sequences thereof, e.g., complements or reverse complements 446). At least some (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 100, 1,000, 10,000, 100,000, 1,000,000, 1,000,000, or more) of the plurality of barcodes 444 can include a cell label 450, a molecular label 448, and a capture sequence 434 (or a complement of the capture sequence, a reverse complement of the capture sequence, or a combination thereof). The molecular labels 448 of at least two barcodes of the plurality of barcodes 444 include different molecular label sequences. At least two barcodes of the plurality of barcodes 444 may include cell labels 450 having the same cell label sequence. The method 400 may include obtaining sequencing data of the plurality of barcoded ssDNA fragments 446 (or complementary sequences thereof); and determining information about the dsDNA (e.g., gDNA) based on the sequences of the plurality of ssDNA fragments 445 in the obtained sequencing data.

[0183] In some embodiments, the dsDNA (e.g., gDNA) is inside the nucleus 452. Method 400 may optionally include permeabilizing the nucleus 452 (e.g., at block 402) to produce a permeabilized nucleus. Method 400 may optionally include fixing the cell containing the nucleus 452 (e.g., at block 402) before permeabilizing the nucleus.

[0184] In some embodiments, method 400 includes denaturing 414 a plurality of nucleic acid fragments 440 to generate a plurality of ssDNA fragments 442. Barcoding 424 the plurality of nucleic acid fragments may include barcoding 424 the plurality of ssDNA fragments 442 with the plurality of barcodes 444 to generate a plurality of barcoded ssDNA fragments 446 and / or their complementary sequences.

[0185] In some embodiments, for any of the methods of sample analysis described herein, the method further comprises contacting cells with a nucleic acid reagent described herein. The nucleic acid reagent may comprise a capture sequence, a barcode, a primer binding site, and a double-stranded DNA binding agent. The cells may be dead cells. The nucleic acid reagent may bind to double-stranded DNA in the dead cells. The method may further comprise washing the cells to remove excess nucleic acid reagent. The method may further comprise lysing the cells, thereby releasing the nucleic acid reagent. The method may further comprise attaching a barcode to the nucleic acid reagent. It is believed that dead cells are permeable to the nucleic acid reagent, while live cells are not permeable or are permeable to only trace amounts of the nucleic acid reagent. Thus, it is believed that the methods described herein can identify dead and live cells by ascertaining whether a nucleic acid reagent is bound to the DNA of a cell (e.g., by determining whether a barcode associated with the nucleic acid reagent is associated with the cell) and / or by determining whether at least a threshold number of nucleic acid reagents are bound to a cell (e.g., by determining whether at least a threshold number of barcodes, e.g., at least 10, 50, 100, 500, 1000, 5000, or 10,000 different barcodes, are associated with the nucleic acid reagent associated with the cell).

[0186] Use of transposomes to generate DNA fragments In some embodiments of the method and kit, DNA fragments can be generated using transposomes. As used herein, a "transposome" comprises (i) a double-stranded nuclease configured to induce double-stranded DNA breaks in a structure containing dsDNA, and (ii) at least two copies of an adapter containing a capture sequence. The adapter can be configured for attachment to the end of the dsDNA. Thus, the adapter can be configured to add a capture sequence to the end of the dsDNA after the corresponding portion induces a double-stranded break in the dsDNA. The double-stranded nuclease can include enzymes such as transposases (e.g., mariner transposases such as Tn5, Tn7, Tn10, Tc3, or Mos1), restriction endonucleases (e.g., EcoRI, NotI, HindIII, HhaI, BamH1, or Sal I), CRISPR-associated proteins (e.g., Cas9 or Cas12a), double-strand-specific nucleases (DSNs), or combinations thereof. It is believed that some double-stranded nucleases, such as transposases, can facilitate the addition of adapters to the ends of dsDNA fragments, while others, such as restriction endonucleases, do not. Thus, transposomes can optionally include a ligase (e.g., T4, T7, or Taq DNA ligase). It is further believed that transposomes can be targeted to specific structures containing dsDNA, such as chromatin, methylated dsDNA, transcription initiation complexes, and the like. Thus, by targeting adapters to structures containing dsDNA, fragmenting the dsDNA, and barcoding the dsDNA to obtain sequence information about the dsDNA, transposomes can provide information about the DNA sequence associated with the structures containing dsDNA. Thus, the transposome may further comprise a moiety that targets the transposome to dsDNA, such as an antibody (e.g., antibody HTA28 or that specifically binds to histone phosphorylation S28 of histone H3) or a fragment thereof, an aptamer (nucleic acid or peptide), or a structure comprising a DNA-binding domain (e.g., a zinc finger binding domain).In any of the methods of sample analysis described herein, the transposome may target a specific structure containing dsDNA, such as chromatin, a specific DNA methylation state, DNA in a specific organelle, etc. It is contemplated that the sample analysis method may identify a specific DNA sequence associated with the structure targeted by the transposome, such as chromatin-accessible DNA, construct DNA, organelle DNA, etc. In some embodiments, a kit for sample analysis is described. The kit may include a transposome described herein and a plurality of barcodes described herein. The barcodes may be immobilized on the particles described herein.

[0187] As an example, generating a plurality of nucleic acid fragments can include contacting dsDNA (e.g., gDNA) with a transposome 428, where the transposome 428 includes a double-stranded nuclease (e.g., transposase) 430 configured to induce double-stranded DNA breaks in a structure comprising the dsDNA and two copies 434a, 434b of an adaptor comprising a capture sequence (e.g., a poly(dT) sequence) to generate a plurality of double-stranded DNA (dsDNA) fragments 440 comprising sequences 434a', 434b' complementary to the capture sequences 434a, 434b, respectively. For example, the adaptor may not include a 5' overhang, such as poly(dT) overhangs 434a, 434b. The double-stranded nuclease (e.g., transposase) 430 can be loaded with two copies 434a, 434b of the adaptor. In some embodiments, the capture sequences 434a, 434b include a poly(dT) region. The sequences 434a', 434b' complementary to the capture sequences may include a poly(dA) region.

[0188] Generating a plurality of nucleic acid fragments can include contacting 410 a dsDNA (e.g., gDNA) with a transposome 428, where the transposome 428 includes a double-stranded nuclease (e.g., transposase) 430 configured to induce double-stranded DNA breaks in a structure comprising the dsDNA and two copies 432a, 432b of an adaptor having 5' overhangs 434a, 434b that include a capture sequence, to generate a plurality of double-stranded DNA (dsDNA) fragments 438, each having two copies 434a, 434b of the 5' overhang. The double-stranded nuclease 430 can be loaded with two copies 432a, 432b of the adaptor. In some embodiments, method 400 may include contacting 412 a plurality of dsDNA fragments 438 having 5' overhangs 434a, 434b with a polymerase to generate a plurality of nucleic acid fragments 440, the plurality of nucleic acid fragments 440 including a plurality of dsDNA fragments each including a sequence 434a', 434b' complementary to at least a portion of the 5' overhang (e.g., a complement or a reverse complement). In some embodiments, none of the plurality of dsDNA fragments 442 includes an overhang (e.g., a 3' overhang or a 5' overhang such as 5' overhangs 434a', 434b').

[0189] Higher signal strength To capture genomic and chromatin accessibility information, the signal, e.g., the number of dsDNA fragments of interest, such as dsDNA fragments for chromatin accessibility analysis, can be further amplified by incorporating a promoter (e.g., a T7 promoter) before the poly(dA) tail of the transposome 428. For example, dsDNA (e.g., gDNA) can be further amplified (e.g., 1000-fold) by incorporating in vitro transcription within the nucleus 452 or cells prior to loading into the single cell line or platform 416. For example, a T7 promoter 502 in the sequence can be added to the end of the dsDNA (e.g., gDNA) fragment.

[0190] After transposition and addition of the poly(dT) sequence and promoter, the fixed cells or nuclei are incubated with an in vitro transcription (IVT) reaction mix. Thousands of copies of RNA containing dsDNA (e.g., gDNA) sequences are generated and contained within the fixed cells or nuclei. Single-cell capture and lysis (e.g., at block 418 in Figures 4A-4B) can be performed as described herein.

[0191] 5A-5B schematically illustrate a non-limiting exemplary method for capturing genomic and chromatin accessibility information from a single cell with improved signal intensity. In some embodiments, adapters 432a, 432b optionally include a promoter sequence. The promoter sequence may include a T7 promoter sequence 502. Generating a plurality of nucleic acid fragments may include transcribing a plurality of dsDNA fragments using in vitro transcription to generate a plurality of ribonucleic acid (RNA) molecules 504. Barcoding 424 the plurality of nucleic acid fragments includes barcoding the plurality of RNA molecules 504.

[0192] Use of restriction enzymes to generate blunt-ended dsDNA fragments In some embodiments, generating a plurality of nucleic acid fragments comprises fragmenting dsDNA (e.g., gDNA) with a restriction enzyme to generate a plurality of dsDNA fragments with blunt ends. Fragmenting dsDNA (e.g., gDNA) may comprise contacting dsDNA (e.g., gDNA) with a restriction enzyme to generate a plurality of dsDNA fragments, each having a blunt end. At least one of the plurality of dsDNA fragments may comprise a blunt end. At least one of the plurality of dsDNA fragments may comprise a 5' overhang or a 3' overhang. None of the plurality of dsDNA fragments may comprise a blunt end. Fragmenting dsDNA (e.g., gDNA) may comprise contacting double-stranded gDNA with a restriction enzyme to generate a plurality of dsDNA fragments with blunt ends. At least one, some, or all of the dsDNA fragments may comprise a blunt end.

[0193] Use of CRISPR-associated proteins to generate dsDNA fragments In some embodiments, generating a plurality of nucleic acid fragments comprises fragmenting dsDNA (e.g., gDNA) using a CRISPR-associated protein, such as Cas9 of Cas12a, to generate a plurality of double-stranded deoxyribonucleic acid (dsDNA) fragments. Fragmenting dsDNA (e.g., gDNA) may comprise contacting the double-stranded gDNA with a CRISPR-associated protein to generate a plurality of dsDNA fragments. At least one, some, or all of the dsDNA fragments may comprise blunt ends. In some embodiments, cleavage of the dsDNA may be targeted to a specific sequence or motif using a guide RNA (gRNA) that targets the specific sequence or motif, whereby the CRISPR-associated protein induces a double-stranded break at the specific sequence or motif.

[0194] Generation of nucleic acid fragments In some embodiments, generating the plurality of nucleic acid fragments (e.g., using a restriction enzyme or a CRISPR-associated protein) includes adding two copies of an adaptor comprising a sequence complementary to the capture sequence to at least a portion (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 100, 1000, 10000, 100000, 1000000, 10000000, or more) of the plurality of dsDNA fragments to generate the plurality of dsDNA fragments (e.g., a plurality of dsDNA fragments with blunt ends) (e.g., block 410 described with reference to Figures 4A-4B). Adding two copies of the adaptor can include ligating two copies of the adaptor to at least a portion (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 100, 1000, 10000, 100000, 1000000, 10000000, or more) of the plurality of dsDNA fragments to generate the plurality of dsDNA fragments comprising the adaptor.

[0195] Use of restriction enzymes to generate dsDNA fragments with overhangs In some embodiments, generating a plurality of nucleic acid fragments includes fragmenting a dsDNA (e.g., gDNA) with a restriction enzyme to generate a plurality of dsDNA fragments with overhangs, thereby eliminating the need to add adapters. Fragmenting a dsDNA (e.g., gDNA) can include contacting the dsDNA (e.g., gDNA) with a restriction enzyme to generate a plurality of dsDNA fragments, wherein at least one of the plurality of dsDNA fragments includes a capture sequence. The capture sequence can be complementary to the sequence of the 5' overhang. The sequence complementary to the capture sequence can include the sequence of the 5' overhang.

[0196] chromatin accessibility 4A-4B, to capture chromatin accessibility information 406a, nuclei 452 can be incubated with enzyme cutters (e.g., transposases, restriction enzymes, and Cas9) and adapters 432a, 432b can be added to dsDNA (e.g., gDNA) fragments. Cleavage can occur at locations where chromatin is exposed (e.g., most exposed, more than average exposed, and desired exposed). For example, transposase 432 can insert adapters 432a, 432b into dsDNA (e.g., gDNA).

[0197] In some embodiments, determining information about the dsDNA (e.g., gDNA) includes determining chromatin accessibility 406a of the dsDNA (e.g., gDNA) based on the sequence and / or abundance of a plurality of ssDNA fragments 442 in the obtained sequencing data. Determining chromatin accessibility 442 of the dsDNA (e.g., gDNA) may include aligning the sequences of the plurality of ssDNA fragments 442 with a reference sequence of dsDNA (e.g., gDNA); and determining whether regions of the dsDNA (e.g., gDNA) corresponding to the ends of the ssDNA fragments of the plurality of ssDNA fragments 442 are accessible or have a certain accessibility (e.g., highly accessible, above average accessibility, and above a threshold or desired degree of accessibility). Determining the chromatin accessibility of dsDNA (e.g., gDNA) may include aligning sequences of a plurality of ssDNA fragments with a reference sequence of dsDNA (e.g., gDNA); and determining the accessibility of regions of the dsDNA (e.g., gDNA) corresponding to ends of ssDNA fragments of the plurality of ssDNA fragments based on the number of ssDNA fragments of the plurality of ssDNA fragments in the sequencing data.

[0198] For example, cleavage may occur at positions where chromatin has above-average accessibility. Regions of dsDNA (e.g., gDNA) corresponding to the ends of ssDNA fragments may have above-average accessibility. Such regions of dsDNA (e.g., gDNA) may have above-average abundance in the resulting sequencing data. As another example, a dsDNA (e.g., gDNA) includes region A-region B1-region B2-region C. If region B1 and region B2 have above-average accessibility while region A and region C have below-average accessibility, region B1 and region B2 may be cleaved (e.g., between region B1 and region B2), while region A and region C are not cleaved. The sequencing data may include above-average abundance of sequences in region B1 and region B2 where cleavage occurs (and near which cleavage occurs). The sequences in region A and region C may be absent (or have low abundance) in the sequencing data. Thus, the chromatin accessibility of dsDNA (e.g., gDNA) can be determined based on the sequence and number of each of multiple ssDNA fragments.

[0199] Genome information To capture genomic information 406b, the nuclei 452 may first be exposed to a reagent (408) to digest nucleosomal structures (e.g., to remove nucleosome / histone proteins) before being subjected to an enzymatic cutter and adding adapters. In some embodiments, determining information about the dsDNA (e.g., gDNA) includes determining genomic information 406b of the dsDNA (e.g., gDNA) based on the sequences of a plurality of ssDNA fragments 442 in the obtained sequencing data. The method may include digesting 408 nucleosomes associated with the double-stranded dsDNA (e.g., gDNA). Determining genomic information of the dsDNA (e.g., gDNA) may include determining at least a partial sequence of the dsDNA (e.g., gDNA) by aligning the sequences of the plurality of ssDNA fragments 442 with a reference sequence of the dsDNA (e.g., gDNA). In some embodiments, the entire or partial genome of the cell may be determined. In some embodiments, the dsDNA is genomic DNA (gDNA) of the cell. In some embodiments, the dsDNA is genomic DNA of an organelle, such as a mitochondrion or chloroplast.

[0200] Methylome Information To capture methylome information 406c, dsDNA (e.g., gDNA) fragments are captured by capture probes 444 and retained single stranded 442, after which methylcytosine bases 454mc are converted to thymine bases using bisulfite treatment 422. The dsDNA (e.g., gDNA) can then be copied by RT 424 or a DNA polymerase.

[0201] In some embodiments, determining information about the dsDNA (e.g., gDNA) includes determining methylome information 406c of the dsDNA (e.g., gDNA) based on sequences of a plurality of ssDNA fragments 442 in the obtained sequencing data. The method may include digesting 408 nucleosomes associated with the dsDNA (e.g., gDNA). Method 400 may include performing bisulfite conversion 422 of cytosine bases of a plurality of single-stranded DNAs 442 to generate a plurality of bisulfite-converted ssDNAs 442b having uracil bases 454u. Barcoding 424 the plurality of ssDNA fragments 442 may include barcoding 424 the plurality of bisulfite-converted ssDNAs 452b with the plurality of barcodes 444 to generate a plurality of barcoded ssDNA fragments 446 and / or their complementary sequences. Determining the methylome information 406c may include determining that positions of a plurality of ssDNA fragments 442 in the sequencing data have a thymine base (or a uracil base 454u) and that corresponding positions in a reference sequence of dsDNA (e.g., gDNA) have a cytosine base, thereby determining that corresponding positions in the dsDNA (e.g., gDNA) have a methylcytosine base 454mc.

[0202] In some embodiments, determining the methylome information comprises a method of sample analysis comprising contacting double-stranded deoxyribonucleic acid (dsDNA) from a cell with a transposome comprising a double-stranded nuclease configured to induce double-stranded DNA breaks in a structure comprising dsDNA loaded with two copies of an adapter having a 5' overhang comprising a capture sequence, to generate a plurality of overhanging dsDNA fragments, each comprising two copies of the 5' overhang. The method may further include contacting the plurality of overhanging dsDNA fragments with a polymerase to generate a plurality of complementary dsDNA fragments, each comprising a sequence complementary to at least a portion of a respective 5' overhang; denaturing the plurality of complementary dsDNA fragments to generate a plurality of single-stranded DNA (ssDNA) fragments; barcoding the plurality of ssDNA fragments using a plurality of barcodes to generate a plurality of barcoded ssDNA fragments, each of the plurality of barcodes comprising a cell label sequence, a molecular label sequence, and a capture sequence, wherein at least two of the plurality of barcodes comprise different molecular label sequences and at least two of the plurality of barcodes comprise the same cell label sequence; obtaining sequencing data of the plurality of barcoded ssDNA fragments; and determining information about the dsDNA based on the sequences of the plurality of barcoded ssDNA fragments in the sequencing data. In some embodiments, the method further includes capturing ssDNA fragments of the plurality of barcoded ssDNA fragments on particles comprising oligonucleotides comprising the capture sequence, the cell label sequence, and the molecular label sequence. For example, the capture sequence may include a poly(dT) sequence that binds to a poly(A) tail on the ssDNA fragment. The captured ssDNA fragment may include a methylated cytidine, which binds to a poly(A) tail on the ssDNA fragment. BisulfiteThe method may include performing a conversion reaction to convert methylated cytidine to thymidine, extending the ssDNA fragment in the 5'-3' direction to generate a barcoded ssDNA fragment containing thymidine, wherein the barcoded ssDNA includes a capture sequence, a molecular beacon sequence, and a cell beacon sequence, extending an oligonucleotide in the 5'-3' direction using a reverse transcriptase or a polymerase, or a combination thereof, to generate a complementary DNA strand complementary to the barcoded ssDNA containing thymidine, denaturing the barcoded ssDNA and the complementary DNA strand to generate a single-stranded sequence, and amplifying the single-stranded sequence.

[0203] In some embodiments, obtaining the methylome information comprises determining that positions of a plurality of ssDNA fragments in the sequencing data have a thymine base and that corresponding positions in a reference sequence of dsDNA have a cytosine base, and determining the methylome information of the plurality of ssDNA fragments. Bisulfite The conversion includes converting methylated cytosine bases to thymine bases and determining that the corresponding positions of the thymine bases in the reference sequence are cytosine bases.

[0204] Multi-omics In some embodiments, the method can include barcoding a plurality of targets (e.g., targets in nuclei 452) using a plurality of barcodes 444 to generate a plurality of barcoded targets; and obtaining sequencing data for the barcoded targets. The targets can be nucleic acid targets, such as mRNA targets, sample-indexing oligonucleotides (e.g., as described in U.S. Patent Application No. 15 / 937,713 (published as U.S. Patent Application Publication No. 2018 / 0346970), which is incorporated herein by reference in its entirety), and oligonucleotides for determining protein expression (e.g., as described in U.S. Patent Application No. 15 / 715028 (published as U.S. Patent Application Publication No. 2018 / 0088112), which is incorporated herein by reference in its entirety). In some embodiments, two or more of genomic, chromatin accessibility, methylome, transcriptome, and proteome information can be determined within a single cell.

[0205] In some embodiments, the method of analyzing a sample includes contacting double-stranded deoxyribonucleic acid (dsDNA) from a cell with a transposome comprising a double-stranded nuclease configured to induce double-stranded DNA cleavage in a structure comprising dsDNA loaded with two copies of an adapter having 5' overhangs comprising a capture sequence to generate a plurality of overhanging dsDNA fragments, each fragment comprising two copies of the 5' overhang; contacting the plurality of overhanging dsDNA fragments with a polymerase to generate a plurality of complementary dsDNA fragments, each fragment comprising a sequence complementary to at least a portion of a respective one of the 5' overhangs; and The method includes denaturing the NA fragments to generate a plurality of single-stranded DNA (ssDNA) fragments, barcoding the plurality of ssDNA fragments using a plurality of barcodes to generate a plurality of barcoded ssDNA fragments, each of the plurality of barcodes comprising a cell label sequence, a molecular label sequence, and a capture sequence, at least two of the plurality of barcodes comprising different molecular label sequences, and at least two of the plurality of barcodes comprising the same cell label sequence, obtaining sequencing data of the plurality of barcoded ssDNA fragments, and determining information about the dsDNA based on the sequences of the plurality of barcoded ssDNA fragments in the sequencing data. In some embodiments, the method further includes contacting a cell with a nucleic acid reagent, the nucleic acid reagent comprising a capture sequence, a barcode, a primer binding site, and a double-stranded DNA binding agent, wherein the cell is a dead cell and the nucleic acid binding reagent binds to the double-stranded DNA in the dead cell, washing the dead cell to remove excess nucleic acid reagent, lysing the dead cell to thereby release the nucleic acid reagent, and barcoding the nucleic acid reagent.

[0206] In some embodiments of the method of sample analysis, the cells are associated with a solid support comprising an oligonucleotide comprising a cell labeling sequence, and the barcoding step comprises barcoding a nucleic acid reagent with the cell labeling sequence.

[0207] In some embodiments of the method of analyzing a sample, the solid support comprises a plurality of oligonucleotides, each of which comprises a cell label sequence and a different molecular label sequence.

[0208] In some embodiments, the method of sample analysis further comprises sequencing the barcoded nucleic acid reagent and determining the presence of dead cells based on the presence of the barcode on the nucleic acid reagent.

[0209] In some embodiments, the method of sample analysis further comprises associating two or more cells with different solid supports each comprising a different cell marker, whereby each of the two or more cells is associated one-to-one with the different cell marker.

[0210] In some embodiments, the method of sample analysis further comprises determining the number of dead cells in the sample based on the number of unique cell markers associated with the barcodes of the nucleic acid reagents.

[0211] In some embodiments, the method of analyzing a sample includes determining the number of molecular label sequences and control barcode sequences having unique sequences associated with the cellular label, including, for each cellular label in the sequencing data, determining the number of molecular label sequences and control barcode sequences having the highest number of unique sequences associated with the cellular label.

[0212] In some embodiments of the methods of sample analysis, the cells are live cells and the nucleic acid reagents do not enter the live cells and therefore do not bind to double-stranded DNA within the live cells.

[0213] In some embodiments, the method of sample analysis further comprises contacting the dead cells with a protein binding reagent associated with a unique identifier oligonucleotide, whereby the protein binding reagent binds to a protein of the dead cells, and barcoding the unique identifier oligonucleotide.

[0214] In some embodiments of the method of sample analysis, the protein binding reagent comprises an antibody, a tetramer, an aptamer, a protein scaffold, an invasin, or a combination thereof. In some embodiments, the protein binding reagent comprises an antibody or fragment thereof, an aptamer, a small molecule, a ligand, a peptide, an oligonucleotide, or any combination thereof. For example, the protein binding reagent can comprise, consist essentially of, or consist of a polyclonal antibody, a monoclonal antibody, a recombinant antibody, a single-chain antibody (scAb), or a fragment thereof, such as Fab, Fv, or scFv. For example, the antibody can comprise, consist essentially of, or consist of an Abseq antibody (see Shahi et al. (2017), Sci Rep. 7:44447, the entire contents of which are incorporated herein by reference). The unique identifier of the protein binding reagent can comprise a nucleotide sequence. In some embodiments, the unique identifier comprises a nucleotide sequence 25-45 nucleotides in length. In some embodiments, the unique identifier does not match the genomic sequence of the sample or cell. In some embodiments, the protein binding reagent may be covalently associated with the unique identifier oligonucleotide. In some embodiments, the protein binding reagent may be covalently associated with the unique identifier oligonucleotide. For example, the protein binding reagent may be associated with the unique identifier oligonucleotide via a linker. In some embodiments, the linker may include a chemical group that reversibly links the oligonucleotide to the protein binding reagent. The chemical group may be conjugated to the linker, for example, via an amine group. In some embodiments, the linker may include a chemical group that forms a stable bond with another chemical group conjugated to the protein binding reagent. For example, the chemical group may be a UV photocleavable group, streptavidin, biotin, an amine, or the like. In some embodiments, the chemical group may be conjugated to the protein binding reagent via a primary amine on an amino acid such as lysine or on the N-terminus. The oligonucleotide may be conjugated to any suitable site on the protein binding reagent, so long as it does not interfere with specific binding between the protein binding reagent and its protein target.In embodiments where the protein binding reagent is an antibody, the oligonucleotide may be located anywhere other than the antigen binding site, for example, in the Fc region, C. H 1 domain, C H 2 domains, C H 3 domains, C L In some embodiments, each protein-binding reagent may be conjugated to an antibody at a domain, etc. In some embodiments, each protein-binding reagent may be conjugated to a single oligonucleotide molecule. In some embodiments, each protein-binding reagent may be conjugated to two or more oligonucleotide molecules, e.g., at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 1,000 or more oligonucleotide molecules, wherein each of the oligonucleotide molecules comprises the same unique identifier.

[0215] In some embodiments of the method of sample analysis, the protein target of the protein-binding reagent is selected from a group comprising 10 to 100 different protein targets, or the cellular component target of the cellular component-binding reagent is selected from a group comprising 10 to 100 different cellular component targets.

[0216] In some embodiments of the method of sample analysis, the protein target of the protein binding reagent comprises a carbohydrate, a lipid, a protein, an extracellular protein, a cell surface protein, a cell marker, a B cell receptor, a T cell receptor, a major histocompatibility complex, a tumor antigen, a receptor, an integrin, an intracellular protein, or any combination thereof.

[0217] In some embodiments of the method of sample analysis, the protein binding reagent comprises an antibody or fragment thereof that binds to a cell surface protein.

[0218] In some embodiments of the method of analyzing a sample, the barcoding is with a barcode comprising a molecular beacon sequence.

[0219] In some embodiments, a method of sample analysis includes contacting dead cells of a sample with a nucleic acid reagent. The nucleic acid reagent can comprise, consist essentially of, or consist of any of the nucleic acid agents described herein. For example, the nucleic acid binding agent can include a capture sequence, a barcode, a primer binding site, and a double-stranded DNA binding agent. By way of example, the barcode can include a cell label, a molecular label, and a target binding region described herein. The nucleic acid reagent can bind to double-stranded DNA in the dead cells. The method can include washing excess nucleic acid reagent from the dead cells, for example, by centrifuging the sample, aspirating fluid from the sample, and applying new fluid, such as a buffer, to the sample. The washing step can remove unbound nucleic acid reagent, while double-stranded DNA-bound nucleic acid binding reagent can remain bound to the double-stranded DNA of the dead cells. For live cells, the washing step is considered to remove all (or all but trace amounts of nucleic acid reagent). The method can include lysing the dead cells. The lysing step can release the nucleic acid reagent from the dead cells. By way of example, dead cells can be lysed by the addition of a cell lysis buffer comprising a detergent (e.g., SDS, Li dodecyl sulfate, Triton X-100, Tween-20, or NP-40), an organic solvent (e.g., methanol or acetone), a digestive enzyme (e.g., proteinase K, pepsin, or trypsin), or any combination thereof. The method can include barcoding the nucleic acid reagent described herein. The barcoding can generate a nucleic acid comprising the barcode of the nucleic acid reagent (or its complement) labeled with the cell label. Optionally, the nucleic acid can further comprise a molecular label. The cell label can associate the nucleic acid reagent with a cell (e.g., a dead cell) on a one-to-one basis, and it is contemplated that the molecular label can be used to quantify the number of nucleic acid reagents associated with a single cell (e.g., a dead cell).

[0220] In some embodiments of the method of sample analysis, barcoding comprises capturing the dead cells on a solid support, such as a bead, wherein the solid support comprises a cell label sequence and a molecular label sequence.

[0221] In some embodiments, the sample analysis method further includes determining the number of unique molecular label sequences associated with each cell label sequence and determining the number of dead cells in the sample based on the number of unique cell label sequences associated with the molecular label sequences. For example, in some embodiments, the presence of a barcode of a nucleic acid reagent described herein can indicate that the cell is a dead cell. For example, in some embodiments, an amount of barcodes of a nucleic acid reagent above a threshold can indicate that the cell is a dead cell. The threshold can include, for example, a limit of detection or an amount of barcodes of a nucleic acid reagent above a negative control, e.g., a known live cell. In some embodiments, an amount of barcodes of at least 10, 50, 100, 500, 1000, 5000, or 10,000 of a nucleic acid reagent associated with a cell can indicate that the cell is a dead cell.

[0222] In some embodiments of the method of sample analysis, determining the number of molecular label sequences and control barcode sequences having unique sequences associated with the cell label comprises determining, for each cell label in the sequencing data, the number of molecular label sequences having the highest number of unique sequences associated with the cell label.

[0223] In some embodiments, the method of sample analysis further comprises contacting the dead cells with a protein-binding reagent associated with a unique identifier oligonucleotide, whereby the protein-binding reagent binds to a protein of the dead cells. The method may further comprise barcoding the unique identifier oligonucleotide. Optionally, the protein-binding reagent may be contacted with the dead cells before washing the dead cells. In some embodiments, the dead cells are contacted with two or more different protein-binding reagents, each associated with a unique identifier oligonucleotide. Thus, if present, at least two different proteins of the dead cells may be bound by different protein-binding reagents.

[0224] In some embodiments of the methods of sample analysis, the protein binding reagent is associated with two or more sample-indexing oligonucleotides having the same sequence.

[0225] In some embodiments of the methods of sample analysis, the protein binding reagent is associated with two or more sample-indexing oligonucleotides having different sample-indexing sequences.

[0226] In some embodiments of the method of sample analysis, the protein binding reagent comprises an antibody, a tetramer, an aptamer, a protein scaffold, an invasin, or a combination thereof.

[0227] In some embodiments of the method of sample analysis, the protein target of the protein-binding reagent is selected from a group comprising 10 to 100 different protein targets, or the cellular component target of the cellular component-binding reagent is selected from a group comprising 10 to 100 different cellular component targets.

[0228] In some embodiments of the method of sample analysis, the protein target of the protein binding reagent comprises a carbohydrate, a lipid, a protein, an extracellular protein, a cell surface protein, a cell marker, a B cell receptor, a T cell receptor, a major histocompatibility complex, a tumor antigen, a receptor, an integrin, an intracellular protein, or any combination thereof.

[0229] In some embodiments of the method of sample analysis, the protein binding reagent comprises an antibody or fragment thereof that binds to a cell surface protein.

[0230] In some embodiments of the methods of sample analysis, the capture sequence and the sequence complementary to the capture sequence are a specific pair of complementary nucleic acids at least 5 nucleotides to about 25 nucleotides in length.

[0231] In some embodiments, a method for analyzing a sample includes contacting double-stranded deoxyribonucleic acid (dsDNA) from a cell with a transposome. The transposome may include a double-stranded nuclease configured to induce double-stranded DNA cleavage in a structure including dsDNA loaded with two copies of an adapter having 5' overhangs comprising a capture sequence to generate a plurality of overhanging dsDNA fragments, each of which comprises two copies of the 5' overhang. The method may include contacting the plurality of overhanging dsDNA fragments with a polymerase to generate a plurality of complementary dsDNA fragments, each of which comprises a sequence complementary to at least a portion of each of the 5' overhangs. The method may include denaturing the plurality of complementary dsDNA fragments to generate a plurality of single-stranded DNA (ssDNA) fragments. The method may include barcoding the plurality of ssDNA fragments with a plurality of barcodes to generate a plurality of barcoded ssDNA fragments, each of the plurality of barcodes comprising a cell label sequence, a molecular label sequence, and a capture sequence. All of the cell labeling sequences associated with a single cell can be the same, so that each single cell is associated with a cell labeling sequence in a one-to-one relationship. At least two of the plurality of barcodes can contain different molecular labeling sequences. The method can include obtaining sequencing data for the plurality of barcoded ssDNA fragments. The method can include quantifying the amount of dsDNA in the cell based on the amount of unique molecular labeling sequences associated with the same cell labeling sequence.

[0232] In some embodiments, the method for analyzing a sample further includes capturing ssDNA fragments of the plurality of ssDNA fragments on a solid support comprising oligonucleotides comprising a capture sequence, a cell label sequence, and a molecular label sequence. The capture sequence may comprise a target binding sequence that hybridizes to a sequence of the ssDNA fragment that is complementary to the target binding sequence. For example, the capture sequence may comprise a poly(dT) sequence that binds to a poly(A) tail on the ssDNA fragment. The method may include extending the ssDNA fragment in the 5'-3' direction to generate a barcoded ssDNA fragment. For example, the extending step may be performed using a DNA polymerase. The barcoded ssDNA may comprise the capture sequence, the molecular label sequence, and the cell label sequence. The method may include extending the oligonucleotide in the 5'-3' direction using a reverse transcriptase or a polymerase, or a combination thereof, to generate a complementary DNA strand complementary to the barcoded ssDNA. The method may include denaturing the barcoded ssDNA and the complementary DNA strand to generate a single-stranded sequence. The method may include amplifying the single-stranded sequence.

[0233] In some embodiments, the sample analysis method further includes bisulfite conversion of cytosine bases of a plurality of ssDNA fragments to generate a plurality of bisulfite-converted ssDNA fragments containing uracil bases. Thus, when a complementary DNA strand complementary to the barcoded ssDNA is generated, the position complementary to the uracil base is expected to contain adenine (rather than guanine, as would be expected if the cytosine base were unmethylated and thus retained the cytosine after the bisulfite conversion process). Therefore, the presence of adenine (rather than guanine) at a position on the complementary DNA strand that is expected to contain guanine may indicate methylation of the cytosine at that position. The presence of adenine can be determined by directly sequencing the complementary DNA strand or by sequencing its complement. Optionally, the sequence can be compared to a reference sequence, such as a genomic reference sequence. The reference sequence can be an electronically stored reference.

[0234] Barcode attachment process In some embodiments, barcoding 424 includes loading cells 416 into a single-cell platform. The ssDNA fragments 442 or nucleic acids may hybridize 420 to a capture sequence 434 for barcoding. The barcoded ssDNA fragments 446, complements, reverse complements 446rc, or combinations thereof may be amplified 426 prior to and / or for sequencing as described with reference to FIG. 3.

[0235] In some embodiments, barcoding 424 can include stochastically barcoding a plurality of ssDNA fragments 442 or a plurality of nucleic acids using a plurality of barcodes 444 to generate a plurality of stochastically barcoded ssDNA fragments 446. Barcoding 424 can include barcoding a plurality of ssDNA fragments 442 using a plurality of barcodes 444 associated with particles 456 to generate a plurality of barcoded ssDNA fragments 446, where the barcodes 444 associated with particles 456 include the same cellular label sequence and at least 100 different molecular label sequences.

[0236] In some embodiments, at least one barcode of the plurality of barcodes may be immobilized on a particle. At least one barcode of the plurality of barcodes may be partially immobilized on a particle. At least one barcode of the plurality of barcodes may be encapsulated in a particle. At least one barcode of the plurality of barcodes may be partially encapsulated in a particle. The particle may be disintegrable (e.g., dissolvable or degradable). The particle may comprise a burstable hydrogel particle. The particle may comprise sepharose beads, streptavidin beads, agarose beads, magnetic beads, conjugated beads, protein A conjugated beads, protein G conjugated beads, protein A / G conjugated beads, protein L conjugated beads, oligo(dT) conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, or any combination thereof. The particles may comprise a material selected from the group consisting of polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic material, ceramic, plastic, glass, methylstyrene, acrylic polymer, titanium, latex, sepharose, cellulose, nylon, silicone, and any combination thereof.

[0237] In some embodiments, the barcode of a particle may include molecular labels having at least 1,000 different molecular label sequences. The barcode of a particle may include molecular labels having at least 10,000 different molecular label sequences. The molecular labels of the barcode may include a random sequence. The particle may include at least 10,000 barcodes.

[0238] Barcoding the plurality of ssDNA fragments may include contacting the plurality of ssDNA fragments with capture sequences of the plurality of barcodes; and transcribing the plurality of ssDNA fragments using the plurality of barcodes to generate a plurality of barcoded ssDNA fragments. The method may include amplifying the plurality of barcoded ssDNA fragments to generate a plurality of amplified barcoded DNA fragments before obtaining sequencing data for the plurality of barcoded ssDNA fragments. Amplifying the plurality of barcoded ssDNA fragments may include amplifying the barcoded ssDNA fragments by polymerase chain reaction (PCR).

[0239] Nucleic Acid Reagents In some embodiments, the nucleic acid reagent comprises, consists essentially of, or consists of a capture sequence, a barcode, a primer binding site, and a double-stranded DNA binder. The barcode of the nucleic acid reagent can include an identifier sequence, indicating that the barcode is associated with the nucleic acid reagent. Optionally, according to the methods and kits described herein, different molecular nucleic acid reagents can include different barcode sequences. The nucleic acids can be used in any of the sample analysis methods described herein. In some embodiments, the kit comprises, consists essentially of, or consists of the nucleic acid reagents described herein. Optionally, the kit further comprises a solid support (e.g., particle) described herein. The multiple barcodes described herein can be immobilized on the solid support.

[0240] An example of a nucleic acid reagent 600 of some embodiments is shown in Figure 6. The nucleic acid reagent 600 may include a double-stranded DNA binder 610. The nucleic acid reagent 600 may include a primer binding site 620, e.g., a PCR handle. The nucleic acid reagent 600 may include a barcode 630. The barcode may include a unique identifier sequence. The nucleic acid reagent 600 may include a capture sequence 640, e.g., a poly(A) tail.

[0241] In some embodiments, the nucleic acid reagent is plasma membrane impermeable. Without being bound by theory, it is believed that such nucleic acid reagents cannot pass through intact cell membranes (or can pass through intact cell membranes in only trace amounts) and therefore do not enter the nucleus of live cells (or do not enter the nucleus of live cells in more than trace amounts). In contrast, the cell membrane of dead cells is not intact, and therefore the nucleic acid reagent can enter the nucleus of dead cells. In some embodiments, the nucleic acid reagent is configured to specifically bind to dead cells, and the nucleic acid reagent does not bind to live cells.

[0242] In some embodiments of the nucleic acid reagent, the capture sequence comprises a poly(A) region.

[0243] In some embodiments of the nucleic acid reagent, the primer binding site comprises a universal primer binding site.

[0244] In some embodiments, a method for binding a nucleic acid reagent to cells is described. The method may include labeling cells of a sample with the nucleic acid reagent. Excess nucleic acid reagent may be washed away. Optionally, the cells are also labeled with a protein-binding reagent, such as an Abseq antibody, associated with one or more barcodes, e.g., unique identifier sequences, described herein. The cells may then be associated with particles having the barcodes immobilized thereon. The nucleic acids (e.g., mRNA) and / or unique identifier sequences (of the protein-binding reagent, such as an Abseq antibody) of the cells and the nuclear binding reagent of the cells may be associated with single-cell labels, e.g., immobilized on a solid support or in a compartment. The nucleic acids may be barcoded with the single-cell labels and molecular labels described herein. A library of barcoded nucleic acids may be created. The library may be sequenced. It is noted that in addition to providing information about the number of proteins and / or nucleic acids of a cell, sequencing can provide information about whether the nucleic acid reagent (or a threshold amount of the nucleic acid reagent) is associated with the cell. Association of a nucleic acid reagent with a cell or association of a threshold amount of a nucleic acid reagent (eg, at least 10, 50, 100, 500, 1000, 5000, or 10000 molecules of the nucleic acid reagent) with a cell can indicate that the cell is a dead cell.

[0245] While various aspects and embodiments are disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and not limitation, with the true scope and spirit being indicated by the following claims.

[0246] Those skilled in the art will appreciate that, for this and other processes and methods disclosed herein, the functions performed in these processes and methods may be performed in differing order. Furthermore, the outlined steps and operations are presented by way of example only, and some of the steps and operations may be optional, combined into fewer steps and operations, or expanded into additional steps and operations, without departing from the essence of the disclosed embodiments.

[0247] In connection with the use of substantially all plural and / or singular terms herein, those skilled in the art can convert from plural to singular and / or from singular to plural where appropriate in the context and / or application. Various singular / plural permutations may be expressly set forth herein for clarity.

[0248] In general, it will be understood by those skilled in the art that the terms used in this specification, and particularly in the appended claims (e.g., the body of the appended claims), are generally intended to be "open" terms (e.g., the term "comprising" should be interpreted as "including, but not limited to," the term "having" should be interpreted as "having at least," the term "including" should be interpreted as "including, but not limited to," etc.). Where a specific number of introductory claim recitations is intended, such intention will be explicitly recited in the claim, and it will be further understood by those skilled in the art that, in the absence of such recitation, no such intention exists. For example, as an aid to understanding, the following appended claims may include the use of the introductory phrases "at least one" and "one or more" to introduce claim recitations. However, the use of such phrases indicates that even when the same claim includes the introductory phrases "one or more" or "at least one" and an indefinite article such as "a" or "an" (e.g., "a" and / or "an" should be interpreted to mean "at least one" or "one or more"), the introduction of a claim recitation with the indefinite article "a" or "an" should not be interpreted as meaning to limit any particular claim containing such an introductory claim recitation to embodiments containing only one such recitation; the same applies to the use of a definite article used to introduce a claim recitation. Moreover, even when a specific number of introductory claim recitations is explicitly recited, those skilled in the art will recognize that such recitation should be interpreted to mean at least the recited number (e.g., a minimum recitation of "two recitations," without any other modifiers, means at least two recitations or more than two recitations).Furthermore, when terms similar to "at least one of A, B, and C, etc." are used, such a configuration is generally intended to have the meaning that one of ordinary skill in the art would understand the term (e.g., "a system having at least one of A, B, and C" would include, but is not limited to, a system having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). When terms similar to "at least one of A, B, or C, etc." are used, such a configuration is generally intended to have the meaning that one of ordinary skill in the art would understand the term (e.g., "a system having at least one of A, B, or C" would include, but is not limited to, a system having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). It will be further understood by those skilled in the art that virtually all disjunctive words and / or phrases expressing two or more alternative terms, whether in the specification, claims, or drawings, should be understood to contemplate the possibility of including one of the terms, either of the terms, or both terms. For example, the phrase "A or B" will be understood to include the possibilities of "A" or "B" or "A and B."

[0249] Furthermore, when features or aspects of the present disclosure are described in terms of a Markush group, one skilled in the art will recognize that the present disclosure is also described thereby in terms of any individual members or subgroups of the Markush group.

[0250] As will be understood by those of skill in the art, for all purposes, including with respect to the provision of a specification, all ranges disclosed herein encompass all possible subranges and combinations of subranges. It will be readily recognized that any recited range fully describes and allows for the same range to be divided into at least two, three, four, five, ten, etc. As a non-limiting example, each range described herein can be readily divided into a lower third, middle third, and upper third, etc. As will be further understood by those of skill in the art, all terms such as "up to," "at least," etc., are inclusive of the recited number and refer to ranges that can be subsequently divided into the subranges described above. Finally, as will be understood by those of skill in the art, a range includes each individual member. Thus, for example, a group having 1 to 3 cells refers to a group having 1, 2, or 3 cells. Similarly, a group having 1 to 5 cells refers to a group having 1, 2, 3, 4, or 5 cells, etc.

[0251] From the foregoing, it will be understood that various embodiments of the present disclosure have been described herein for purposes of illustration and that various modifications can be made without departing from the scope and spirit of the present disclosure. Accordingly, the various embodiments disclosed herein are not intended to be limiting, with the true scope and spirit being indicated by the following claims. Another aspect of the present invention may be as follows. [1] A method of sample analysis, comprising: contacting double-stranded deoxyribonucleic acid (dsDNA) from the cell with a transposome comprising a double-stranded nuclease configured to induce double-stranded DNA breaks in a structure comprising the dsDNA and two copies of an adapter having a 5' overhang comprising a capture sequence to generate a plurality of overhanging dsDNA fragments, each comprising two copies of the 5' overhang; barcoding the plurality of overhanging DNA fragments using a plurality of barcodes to generate a plurality of barcoded DNA fragments, each of the plurality of barcodes comprising a cell labeling sequence, a molecular labeling sequence, and the capture sequence, at least two of the plurality of barcodes comprising different molecular labeling sequences, and at least two of the plurality of barcodes comprising the same cell labeling sequence; detecting the sequences of the plurality of barcoded DNA fragments; determining information relating the sequence of the dsDNA to the structure comprising the dsDNA based on the sequences of the plurality of barcoded DNA fragments in the sequencing data; A method comprising: [2] contacting the plurality of overhanging dsDNA fragments with a polymerase to produce a plurality of complementary dsDNA fragments, each of which comprises a sequence complementary to at least a portion of the 5' overhang; denaturing the plurality of complementary dsDNA fragments to produce a plurality of single-stranded DNA (ssDNA) fragments; The method of claim 1, further comprising: wherein the ssDNA fragments are barcoded, thereby barcoding the DNA fragments. [3] A method of sample analysis, comprising: generating a plurality of nucleic acid fragments from double-stranded deoxyribonucleic acid (dsDNA) from a cell, each of the plurality of nucleic acid fragments comprising a capture sequence, a complement of the capture sequence, a reverse complement of the capture sequence, or a combination thereof; barcoding the plurality of nucleic acid fragments using a plurality of barcodes to generate a plurality of barcoded DNA fragments, each of the plurality of barcodes comprising a cell label sequence, a molecular label sequence, and the capture sequence, at least two of the plurality of barcodes comprising different molecular label sequences, and at least two of the plurality of barcodes comprising the same cell label sequence; detecting the sequences of the plurality of barcoded DNA fragments; A method comprising: [4] The method according to [3], further comprising a step of determining information relating the sequence of the dsDNA to a structure containing the dsDNA based on the sequences of the plurality of barcoded DNA fragments. [5] The step of generating a plurality of nucleic acid fragments comprises: contacting the dsDNA with a transposome comprising a double-stranded nuclease configured to induce double-stranded DNA breaks in a structure comprising the dsDNA and two copies of an adapter comprising the capture sequence to generate a plurality of complementary dsDNA fragments, each fragment comprising a sequence complementary to the capture sequence. The method according to [3] or [4], comprising: [6] The step of generating a plurality of nucleic acid fragments comprises: contacting the dsDNA with a transposome comprising a double-stranded nuclease configured to induce double-stranded DNA breaks in a structure comprising the dsDNA and two copies of an adapter having a 5' overhang comprising a capture sequence to generate a plurality of overhanging dsDNA fragments, each fragment comprising two copies of the 5' overhang; contacting the plurality of overhanging dsDNA fragments comprising the 5' overhangs with a polymerase to produce a plurality of complementary dsDNA fragments, each of the dsDNA fragments comprising a sequence complementary to at least a portion of the 5' overhang; The method according to [3] or [4], comprising: [7] The method according to any one of [3] to [6], wherein the barcoded DNA fragment is a barcoded single-stranded DNA. [8] The method according to [1] or [6], wherein none of the plurality of complementary dsDNA fragments contains an overhang. [9] The method according to any one of [1], [2], or [5] to [7], wherein the adapter comprises a DNA terminal sequence of a transposon.

[10] The method according to any one of [1] or [5] to [8], wherein the double-stranded nuclease comprises a transposase such as Tn5 transposase.

[11] The method according to any one of [3] to

[10] above, wherein the step of generating a plurality of nucleic acid fragments comprises a step of fragmenting the dsDNA to generate a plurality of dsDNA fragments.

[12] The method according to

[11] , wherein the step of fragmenting the dsDNA comprises contacting the dsDNA with a restriction enzyme to generate the plurality of dsDNA fragments, each having a blunt end.

[13] The method according to

[11] , wherein at least one of the plurality of dsDNA fragments comprises a blunt end.

[14] The method according to

[11] , wherein at least one of the plurality of dsDNA fragments comprises a 5' overhang and / or a 3' overhang.

[15] The method according to

[11] , wherein none of the plurality of dsDNA fragments contains a blunt end.

[16] The method according to

[11] , wherein the step of fragmenting the dsDNA comprises contacting the dsDNA with a CRISPR-associated protein such as Cas9 or Cas12a to generate the plurality of dsDNA fragments.

[17] The step of generating a plurality of nucleic acid fragments comprises: adding two copies of an adaptor comprising a sequence complementary to the capture sequence to the plurality of dsDNA fragments to generate a plurality of nucleic acid fragments. The method according to any one of

[11] to

[16] above, comprising:

[18] The method according to

[17] , wherein the step of adding the two copies of the adapter comprises ligating the two copies of the adapter to the plurality of dsDNA fragments to generate the plurality of nucleic acid fragments containing the adapter.

[19] The method according to any one of [1] to

[18] above, wherein the capture sequence includes a poly(dT) region.

[20] The method according to any one of [2] or [5] to

[18] , wherein the sequence complementary to the capture sequence comprises a poly(dA) region.

[21] The method according to

[11] , wherein the step of fragmenting the dsDNA comprises contacting the dsDNA with a restriction enzyme to generate the plurality of dsDNA fragments, and at least one of the plurality of dsDNA fragments comprises the capture sequence.

[22] The method according to

[14] , wherein the capture sequence is complementary to the sequence of the 5' overhang.

[23] The method according to

[22] , wherein the sequence complementary to the capture sequence includes the sequence of the 5' overhang.

[24] The method according to any one of [1] to

[23] above, wherein the dsDNA is present in the nucleus during the contact.

[25] The method according to

[24] , further comprising the step of permeabilizing the nuclei to produce permeabilized nuclei.

[26] The method described in

[252] 4, comprising a step of fixing cells containing the nuclei before permeabilizing the nuclei.

[27] The method according to any one of [5] to

[26] , comprising a step of denaturing the plurality of nucleic acid fragments to generate a plurality of ssDNA fragments, wherein the step of attaching barcodes to the plurality of nucleic acid fragments comprises a step of attaching barcodes to the plurality of ssDNA fragments using the plurality of barcodes to generate the plurality of barcoded ssDNA fragments.

[28] The method according to any one of [5] to

[27] above, wherein the adapter comprises a promoter sequence.

[29] The method according to

[28] , wherein the step of generating the plurality of nucleic acid fragments comprises a step of transcribing the plurality of dsDNA fragments using in vitro transcription to generate a plurality of ribonucleic acid (RNA) molecules, and the step of attaching barcodes to the plurality of nucleic acid fragments comprises a step of attaching barcodes to the plurality of RNA molecules.

[30] The method described in

[28] or

[29] , wherein the promoter sequence comprises a T7 promoter sequence.

[31] The method according to any one of [1], [2], or [4] to

[30] , wherein the step of determining information relating the sequence of the dsDNA to the structure includes a step of determining the chromatin accessibility of the dsDNA based on the sequences of the plurality of barcoded DNA fragments in the obtained sequencing data.

[32] The step of determining the chromatin accessibility of the dsDNA comprises: aligning the sequences of the plurality of barcoded DNA fragments with a reference sequence of the dsDNA; identifying regions of the dsDNA corresponding to ends of barcoded DNA fragments of the plurality of barcoded DNA fragments as having accessibility above a threshold; The method according to

[31] , comprising:

[33] The step of determining the chromatin accessibility of the dsDNA comprises: aligning the sequences of the plurality of ssDNA fragments with the dsDNA reference sequence; determining the accessibility of regions of the dsDNA corresponding to ends of ssDNA fragments of the plurality of ssDNA fragments based on the number of ssDNA fragments of the plurality of ssDNA fragments in the sequencing data; The method according to

[31] , comprising:

[34] The method described in any one of [1], [2], or [4] to

[30] , wherein the step of determining the information about the dsDNA includes a step of determining genomic information of the dsDNA based on the sequences of the plurality of barcoded DNA fragments in the obtained sequencing data.

[35] The method according to

[34] , further comprising the step of digesting nucleosomes associated with the dsDNA.

[36] The method according to

[34] or

[35] , wherein the step of determining the genomic information of the dsDNA includes a step of determining at least a partial sequence of the dsDNA by aligning the sequences of the plurality of barcoded DNA fragments with a reference sequence of the dsDNA.

[37] The method according to any one of [1], [2], or [4] to

[36] , wherein the step of determining information relating the sequence of the dsDNA to the structure containing the dsDNA comprises a step of determining methylome information of the dsDNA based on the sequences of the plurality of barcoded DNA fragments in the obtained sequencing data.

[38] The method according to

[37] , further comprising the step of digesting nucleosomes associated with the dsDNA of the cell.

[39] The method according to

[37] or

[38] , further comprising a step of bisulfite converting cytosine bases in a plurality of single-stranded DNAs of the plurality of overhanging DNA fragments or the plurality of nucleic acid fragments to generate a plurality of bisulfite-converted ssDNAs containing uracil bases.

[40] The method described in

[39] , wherein the step of attaching barcodes to the plurality of overhanging DNA fragments or the step of attaching barcodes to the plurality of nucleic acid fragments comprises a step of attaching barcodes to a plurality of bisulfite-converted ssDNA fragments using the plurality of barcodes to generate a plurality of barcoded ssDNA fragments.

[41] The step of determining methylome information comprises: determining whether positions of the plurality of barcoded DNA fragments in the sequencing data have a thymine base and whether the corresponding positions in the dsDNA reference sequence have a cytosine base, thereby determining that the corresponding positions in the dsDNA have a methylcytosine base; The method according to any one of

[37] to

[40] above, comprising:

[42] The step of attaching the barcode comprises: using the plurality of barcodes to stochastically barcode the plurality of DNA fragments or the plurality of nucleic acid fragments to generate a plurality of stochastically barcoded ssDNA fragments; The method according to any one of [1] to

[41] above, comprising:

[43] The step of attaching the barcode comprises: barcoding the plurality of DNA fragments or the plurality of nucleic acid fragments using the plurality of barcodes associated with the particles to generate the plurality of barcoded ssDNA fragments. The method according to any one of [1] to

[41] , wherein the barcodes associated with the particles comprise the same cell marker sequence and at least 100 different molecular marker sequences.

[44] The method according to

[43] , wherein at least one barcode of the plurality of barcodes is fixed on the particle.

[45] The method according to

[43] , wherein at least one barcode of the plurality of barcodes is partially immobilized on the particle.

[46] The method according to

[43] , wherein at least one barcode of the plurality of barcodes is encapsulated in the particle.

[47] The method according to

[43] , wherein at least one barcode of the plurality of barcodes is partially encapsulated in the particle.

[48] ​​The method according to any one of

[44] to

[47] , wherein the particles are disintegrable.

[49] The method according to any one of

[44] to

[47] , wherein the particles include disintegrable hydrogel particles.

[50] The method according to any one of

[43] to

[49] , wherein the particles comprise sepharose beads, streptavidin beads, agarose beads, magnetic beads, conjugate beads, protein A conjugate beads, protein G conjugate beads, protein A / G conjugate beads, protein L conjugate beads, oligo(dT) conjugate beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, or any combination thereof.

[51] The method according to any one of

[43] to

[50] , wherein the particles comprise a material selected from the group consisting of polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic material, ceramic, plastic, glass, methylstyrene, acrylic polymer, titanium, latex, Sepharose, cellulose, nylon, silicone, and any combination thereof.

[52] The method described in any one of

[43] to

[51] , wherein the barcode of the particle includes molecular labels having at least 1,000 different molecular label sequences.

[53] The method described in any one of

[43] to

[52] , wherein the barcode of the particle includes molecular labels having at least 10,000 different molecular label sequences.

[54] The method according to any one of

[42] to

[53] , wherein the molecular label of the barcode comprises a random sequence.

[55] The method according to any one of

[43] to

[54] , wherein the particles contain at least 10,000 barcodes.

[56] The step of attaching barcodes to the plurality of ssDNA fragments comprises: contacting the plurality of ssDNA fragments with the capture sequences of the plurality of barcodes; transcribing the plurality of ssDNA fragments using the plurality of barcodes to generate the plurality of barcoded ssDNA fragments; The method according to any one of

[42] to

[55] above, comprising:

[57] The method described in any one of

[42] to

[56] , comprising a step of amplifying the plurality of barcoded ssDNA fragments to generate a plurality of amplified barcoded DNA fragments before obtaining the sequencing data of the plurality of barcoded ssDNA fragments.

[58] The method described in

[57] , wherein the step of amplifying the plurality of barcoded ssDNA fragments includes a step of amplifying the barcoded ssDNA fragments by polymerase chain reaction (PCR).

[59] barcoding a plurality of targets of the nuclei using the plurality of barcodes to generate a plurality of barcoded targets; obtaining sequencing data for said barcoded targets; The method according to any one of [1] to

[58] above, comprising:

[60] The method according to any one of [1] to

[59] , wherein the dsDNA from the cell is selected from the group consisting of nuclear DNA, nucleolar DNA, genomic DNA, mitochondrial DNA, chloroplast DNA, construct DNA, viral DNA, or a combination of two or more of the above-listed items.

[61] The method according to any one of [1], [2], or [4] to

[60] , wherein the 5' overhang comprises a poly(dT) sequence.

[62] Capturing ssDNA fragments of the plurality of barcoded DNA fragments on particles comprising oligonucleotides including the capture sequence, the cell label sequence, and the molecular label sequence, wherein the capture sequence comprises a poly(dT) sequence that binds to a poly(A) tail on the ssDNA fragment, and the captured ssDNA fragment comprises a methylated cytidine; performing a bisulfide conversion reaction on the ssDNA fragments to convert the methylated cytidines to thymidines; extending the ssDNA fragment in a 5'-3' direction to generate the barcoded ssDNA fragment containing the thymidine, wherein the barcoded ssDNA comprises the capture sequence, a molecular beacon sequence, and a cell beacon sequence; extending the oligonucleotide in the 5'-3' direction using a reverse transcriptase or a polymerase, or a combination thereof, to generate a complementary DNA strand complementary to the barcoded ssDNA containing the thymidine; denaturing the barcoded ssDNA and the complementary DNA strand to generate a single-stranded sequence; amplifying the single-stranded sequence; The method according to any one of [2], [7], or [9] to

[61] , further comprising:

[63] The method of

[62] , further comprising the step of determining whether a position of the ssDNA fragment in the sequencing data has a thymine base after the bisulfide conversion reaction and whether the corresponding position in the reference sequence of the dsDNA has a cytosine base, thereby indicating that the position of the ssDNA fragment contains a methylated cytosine.

[64] The method according to any one of [1] or [2] or [5] to

[63] , wherein the double-stranded nuclease of the transposome is selected from the group consisting of a transposase, a restriction endonuclease, a CRISPR-associated protein, a double-strand-specific nuclease, or a combination thereof.

[65] The method described in any one of [1] or [2] or [5] to

[64] , wherein the transposome further comprises a DNA-binding domain that binds to the structure containing an antibody or a fragment thereof, an apatocyte, or dsDNA.

[66] The method according to any one of [1], [2], or [5] to

[65] , wherein the transposome further comprises a ligase.

[67] capture sequence; Barcodes and; a primer binding site; Double-stranded DNA binding agent A nucleic acid reagent comprising:

[68] The nucleic acid reagent described in

[67] above, which is impermeable to the plasma membrane.

[69] The nucleic acid reagent according to

[67] or

[68] , which is configured to specifically bind to dead cells.

[70] The nucleic acid reagent according to

[67] or

[68] above, which does not bind to living cells.

[71] The nucleic acid reagent according to any one of

[67] to

[70] above, wherein the capture sequence includes a poly(A) region.

[72] The nucleic acid reagent according to any one of

[67] to

[71] , wherein the primer binding site comprises a universal primer binding site.

[73] A step of contacting a cell with a nucleic acid reagent, the nucleic acid reagent comprising: capture sequence; Barcode; a primer binding site; and double-stranded DNA binding agents wherein the cells are dead cells and the nucleic acid binding reagent binds to double-stranded DNA in the dead cells; washing the dead cells to remove excess nucleic acid binding reagent; lysing the dead cells, thereby releasing the nucleic acid binding reagent; attaching a barcode to the nucleic acid binding reagent; The method according to any one of [1] to

[66] above, further comprising:

[74] The method of

[73] , wherein the cells are associated with a solid support comprising an oligonucleotide comprising a cell labeling sequence, and the barcoding step comprises barcoding the nucleic acid binding reagent with the cell labeling sequence.

[75] The method according to

[74] , wherein the solid support comprises a plurality of the oligonucleotides, each of which contains the cell label sequence and a different molecular label sequence.

[76] Sequencing the barcoded nucleic acid binding reagent; determining the presence of dead cells based on the presence of the barcode on the nucleic acid reagent; The method according to any one of

[73] to

[75] above, further comprising:

[77] The method according to any one of

[73] to

[76] , further comprising the step of associating two or more cells with different solid supports each containing a different cell marker, thereby associating each of the two or more cells with a different cell marker in a one-to-one manner.

[78] The method described in

[77] , further comprising a step of determining the number of dead cells in the sample based on the number of unique cell markers associated with the barcode of the nucleic acid reagent.

[79] A method according to any one of

[73] to

[78] , wherein the step of determining the number of molecular marker sequences and control barcode sequences having unique sequences associated with the cell marker includes a step of determining, for each cell marker in the sequencing data, the number of molecular marker sequences and control barcode sequences having the highest number of unique sequences associated with the cell marker.

[80] The method according to any one of

[67] to

[71] above, wherein the nucleic acid binding reagent does not enter living cells and therefore does not bind to double-stranded DNA within the living cells.

[81] contacting dead cells with a protein-binding reagent associated with a unique identifier oligonucleotide, whereby the protein-binding reagent binds to a protein in the dead cells; barcoding said unique identifier oligonucleotides; The method according to any one of

[73] to

[80] above, further comprising:

[82] The method of

[81] , wherein the protein binding reagent comprises an antibody, a tetramer, an aptamer, a protein scaffold, an invasin, or a combination thereof.

[83] The method according to

[81] or

[82] , wherein the protein target of the protein binding reagent is selected from a group containing 10 to 100 different protein targets, or the cellular component target of the cellular component binding reagent is selected from a group containing 10 to 100 different cellular component targets.

[84] The method of any one of

[81] to

[83] , wherein the protein target of the protein binding reagent comprises a carbohydrate, a lipid, a protein, an extracellular protein, a cell surface protein, a cell marker, a B cell receptor, a T cell receptor, a major histocompatibility complex, a tumor antigen, a receptor, an integrin, an intracellular protein, or any combination thereof.

[85] The method according to any one of

[81] to

[84] , wherein the protein binding reagent comprises an antibody or a fragment thereof that binds to a cell surface protein.

[86] The method according to any one of

[73] to

[85] , wherein the barcode attachment step is performed using a barcode containing a molecular marker sequence.

[87] A method of sample analysis, comprising: contacting dead cells of the sample with a nucleic acid binding reagent, said nucleic acid binding reagent comprising: capture sequence; Barcode; a primer binding site; and double-stranded DNA binding agents wherein the nucleic acid binding reagent binds to double-stranded DNA in the dead cells; removing excess nucleic acid binding reagent from the dead cells; lysing the dead cells, thereby releasing the nucleic acid binding reagent from the dead cells; attaching a barcode to the nucleic acid binding reagent; A method comprising:

[88] The method described in

[87] , wherein the barcoding step includes capturing the dead cells on a solid support such as a bead, and the solid support comprises a cell labeling sequence and a molecular labeling sequence.

[89] determining the number of unique molecular marker sequences associated with each cellular marker sequence; determining the number of dead cells in the sample based on the number of unique cell-marker sequences associated with the molecular marking sequence; The method according to

[87] or

[88] , further comprising:

[90] The step of determining the number of molecular marker sequences and control barcode sequences having unique sequences associated with the cell marker comprises: determining, for each cell label in the sequencing data, the number of molecular label sequences having the highest number of unique sequences associated with said cell label; The method according to any one of

[87] to

[89] above, comprising:

[91] contacting dead cells with a protein-binding reagent associated with a unique identifier oligonucleotide, whereby the protein-binding reagent binds to a protein in the dead cells; barcoding said unique identifier oligonucleotides; The method according to any one of

[79] to

[90] above, further comprising:

[92] The method described in any one of

[87] to

[91] , wherein the protein binding reagent is associated with two or more sample-indexed oligonucleotides having the same sequence.

[93] The method described in any one of

[87] to

[92] , wherein the protein binding reagent is associated with two or more sample index-added oligonucleotides having different sample index-added sequences.

[94] The method according to any one of

[88] to

[93] , wherein the protein binding reagent comprises an antibody, a tetramer, an aptamer, a protein scaffold, an invasin, or a combination thereof.

[95] The method according to any one of

[88] to

[94] , wherein the protein target of the protein binding reagent is selected from a group containing 10 to 100 different protein targets, or the cellular component target of the cellular component binding reagent is selected from a group containing 10 to 100 different cellular component targets.

[96] The method according to any one of

[88] to

[95] , wherein the protein target of the protein binding reagent comprises a carbohydrate, a lipid, a protein, an extracellular protein, a cell surface protein, a cell marker, a B cell receptor, a T cell receptor, a major histocompatibility complex, a tumor antigen, a receptor, an integrin, an intracellular protein, or any combination thereof.

[97] The method according to any one of

[88] to

[96] , wherein the protein binding reagent comprises an antibody or a fragment thereof that binds to a cell surface protein.

[98] A method of sample analysis, comprising: contacting double-stranded deoxyribonucleic acid (dsDNA) from the cell with a transposome comprising a double-stranded nuclease configured to induce double-stranded DNA breaks in a structure comprising the dsDNA and two copies of an adapter having a 5' overhang comprising a capture sequence to generate a plurality of overhanging dsDNA fragments, each comprising two copies of the 5' overhang; contacting the plurality of overhanging dsDNA fragments with a polymerase to generate a plurality of complementary dsDNA fragments, each fragment comprising a sequence complementary to at least a portion of each of the 5' overhangs; denaturing the plurality of complementary dsDNA fragments to generate a plurality of single-stranded DNA (ssDNA) fragments; barcoding the plurality of ssDNA fragments using a plurality of barcodes to generate a plurality of barcoded ssDNA fragments, each of the plurality of barcodes comprising a cell labeling sequence, a molecular labeling sequence, and the capture sequence, at least two of the plurality of barcodes comprising different molecular labeling sequences, and the plurality of barcodes comprising the same cell labeling sequence; obtaining sequencing data for the plurality of barcoded ssDNA fragments; quantitating the amount of dsDNA in the cells based on the amount of unique molecular landmark sequences associated with the same cellular landmark sequence; A method comprising:

[99] Capturing ssDNA fragments of the plurality of ssDNA fragments on a solid support comprising an oligonucleotide comprising the capture sequence, the cell label sequence, and the molecular label sequence, wherein the capture sequence comprises a poly(dT) sequence that binds to the poly(A) tails on the ssDNA fragments; extending the ssDNA fragment in a 5'-3' direction to generate the barcoded ssDNA fragment, wherein the barcoded ssDNA comprises the capture sequence, a molecular beacon sequence, and a cell beacon sequence; extending the oligonucleotide in the 5'-3' direction using a reverse transcriptase or a polymerase, or a combination thereof, to generate a complementary DNA strand complementary to the barcoded ssDNA; denaturing the barcoded ssDNA and the complementary DNA strand to generate a single-stranded sequence; amplifying the single-stranded sequence; The method of claim 98, further comprising:

[100] The method described in

[100] further comprises a step of bisulfite converting the cytosine bases of the plurality of ssDNA fragments to produce a plurality of bisulfite converted ssDNA fragments containing uracil bases.

[101] The method described in

[60] , wherein the construct DNA is selected from the group consisting of a plasmid, a cloning vector, an expression vector, a hybrid vector, a minicircle, a cosmid, a viral vector, a BAC, a YAC, and a HAC.

[102] The amount of construct DNA in the sample ranges from 1 to approximately 1 × 10 6 The method according to

[101] , wherein the range of the construct DNA is 100.

[103] The dsDNA comprises viral DNA, and the single cell load of viral DNA is about 1 x 10 2 ~1×10 6 The method according to

[60] , wherein the range is

[104] The method according to any one of [1] to

[103] above, which comprises single-cell analysis.

[105] A kit for sample analysis, comprising: A transposome comprising: a double-stranded nuclease configured to induce double-stranded DNA breaks in a structure containing dsDNA; and Two copies of the adapter with a 5' overhang containing the capture sequence and a transposome comprising: a plurality of barcodes, each barcode may include a cell labeling sequence, a molecular labeling sequence, and the capture sequence, at least two of the plurality of barcodes include different molecular labeling sequences, and at least two of the plurality of barcodes include the same cell labeling sequence; Kit including:

[106] The kit described in

[105] , wherein the double-stranded nuclease comprises a transposase, a restriction endonuclease, a CRISPR-associated protein, a double-strand-specific nuclease (DSN), or a combination thereof.

[107] The kit according to

[105] or

[106] , wherein the transposome further comprises a ligase.

[108] The kit described in any one of

[105] to

[107] , wherein the plurality of barcodes includes at least 10, 50, 100, 500, 1000, 5000, 10000, 50000, or 100000 different molecular labels.

[109] The kit according to any one of

[105] to

[108] , wherein the barcode is fixed on a particle.

[110] The kit described in

[109] , wherein all of the barcodes immobilized on each particle contain the same cell-labeling sequence, and different particles contain different cell-labeling sequences.

[111] The kit according to any one of

[105] to

[108] , wherein the barcodes are distributed among the wells of the substrate.

[112] The kit described in

[111] , wherein all of the barcodes divided into each well contain the same cell-marker sequence, and different wells contain different cell-marker sequences.

Claims

1. 1. A method of sample analysis comprising: contacting double-stranded deoxyribonucleic acid (dsDNA) from the cell with a transposome comprising a double-stranded nuclease configured to induce double-stranded DNA breaks in a structure comprising the dsDNA and two copies of an adapter having a 5' overhang comprising a capture sequence to generate a plurality of overhanging dsDNA fragments, each comprising two copies of the 5' overhang; barcoding the plurality of overhanging dsDNA fragments with a plurality of barcodes to generate a plurality of barcoded DNA fragments, each of the plurality of barcodes comprising a cell labeling sequence, a molecular labeling sequence, and the capture sequence, at least two of the plurality of barcodes comprising different molecular labeling sequences, and at least two of the plurality of barcodes comprising the same cell labeling sequence; barcoding a plurality of nuclear targets using the plurality of barcodes to generate a plurality of barcoded targets, wherein the targets are mRNA; obtaining sequencing data for the barcoded targets; detecting the sequences of the plurality of barcoded DNA fragments; determining information relating the sequence of the dsDNA to the structure comprising the dsDNA based on the sequences of the plurality of barcoded DNA fragments in the sequencing data; A method comprising:

2. contacting the plurality of overhanging dsDNA fragments with a polymerase to generate a plurality of complementary dsDNA fragments, each of which comprises a sequence complementary to at least a portion of the 5' overhang; denaturing the plurality of complementary dsDNA fragments to produce a plurality of single-stranded DNA (ssDNA) fragments; wherein the ssDNA fragments are barcoded, thereby barcoding the DNA fragments; The method of claim 1 , wherein the adapter may comprise a DNA terminal sequence of a transposon.

3. 1. A method of sample analysis comprising: generating a plurality of nucleic acid fragments from double-stranded deoxyribonucleic acid (dsDNA) from a cell, each of the plurality of nucleic acid fragments comprising a capture sequence, a complement of the capture sequence, a reverse complement of the capture sequence, or a combination thereof; barcoding the plurality of nucleic acid fragments using a plurality of barcodes to generate a plurality of barcoded DNA fragments, each of the plurality of barcodes comprising a cell label sequence, a molecular label sequence, and the capture sequence, at least two of the plurality of barcodes comprising different molecular label sequences, and at least two of the plurality of barcodes comprising the same cell label sequence; barcoding a plurality of nuclear targets using the plurality of barcodes to generate a plurality of barcoded targets, wherein the targets are mRNA; obtaining sequencing data for the barcoded targets; detecting the sequences of the plurality of barcoded DNA fragments; determining information relating the sequence of the dsDNA to a structure comprising the dsDNA based on the sequences of the plurality of barcoded DNA fragments; A method comprising:

4. The step of generating a plurality of nucleic acid fragments comprises: (a) contacting the dsDNA with a transposome comprising a double-stranded nuclease configured to induce double-stranded DNA breaks in a structure comprising the dsDNA and two copies of an adapter comprising the capture sequence to generate a plurality of complementary dsDNA fragments, each fragment comprising a sequence complementary to the capture sequence; or (b) contacting the dsDNA with a transposome comprising a double-stranded nuclease configured to induce double-stranded DNA breaks in a structure comprising the dsDNA and two copies of an adapter having a 5' overhang comprising a capture sequence to generate a plurality of overhanging dsDNA fragments, each comprising two copies of the 5' overhang; contacting the plurality of overhanging dsDNA fragments comprising the 5' overhangs with a polymerase to produce a plurality of complementary dsDNA fragments, each of the complementary dsDNA fragments comprising a sequence complementary to at least a portion of the 5' overhang; The method of claim 3, comprising:

5. The method of claim 3 or 4, wherein the barcoded DNA fragment is a barcoded single-stranded DNA.

6. 6. The method of claim 2, 4 or 5, wherein none of the plurality of complementary dsDNA fragments comprises an overhang.

7. (i) the adapter comprises a DNA terminal sequence of a transposon, and / or (ii) the double-stranded nuclease comprises a transposase such as Tn5 transposase; The method according to any one of claims 1 or 4 to 6.

8. The method of any one of claims 2 or 4 to 7, wherein the sequence complementary to the capture sequence comprises a poly(dA) region.

9. The step of generating a plurality of nucleic acid fragments comprises: fragmenting the dsDNA to generate a plurality of dsDNA fragments; and optionally (i) fragmenting the dsDNA comprises contacting the dsDNA with a restriction enzyme to generate the plurality of dsDNA fragments, each having a blunt end; and / or (ii) at least one of the plurality of dsDNA fragments comprises a blunt end; and / or (iii) at least one of the plurality of dsDNA fragments comprises a 5' overhang and / or a 3' overhang; and / or (iv) none of the plurality of dsDNA fragments comprises a blunt end; and / or (v) fragmenting the dsDNA comprises contacting the dsDNA with a CRISPR-associated protein, such as Cas9 or Cas12a, to generate the plurality of dsDNA fragments; and / or adding two copies of an adaptor to the plurality of dsDNA fragments, the adaptor comprising a sequence complementary to a capture sequence, to generate a plurality of nucleic acid fragments, and the adding the two copies of the adaptor may comprise ligating the two copies of the adaptor to the plurality of dsDNA fragments to generate the plurality of nucleic acid fragments comprising the adaptors. The method according to any one of claims 3 to 8.

10. determining the information relating the sequence of the dsDNA to the structure, (a) determining chromatin accessibility of the dsDNA based on the sequences of the plurality of barcoded DNA fragments in the obtained sequencing data, wherein the determining chromatin accessibility of the dsDNA includes: (i) aligning the sequences of the plurality of barcoded DNA fragments with the dsDNA reference sequence; and identifying regions of the dsDNA corresponding to ends of barcoded DNA fragments of the plurality of barcoded DNA fragments as having an accessibility above a threshold; or (ii) aligning the sequences of the plurality of ssDNA fragments with the dsDNA reference sequence; and determining the accessibility of regions of the dsDNA corresponding to ends of ssDNA fragments of the plurality of ssDNA fragments based on the number of ssDNA fragments of the plurality of ssDNA fragments in the sequencing data; and and / or (b) determining methylome information of the dsDNA based on the sequences of the plurality of barcoded DNA fragments in the obtained sequencing data; (i) the following steps: (A) digesting nucleosomes associated with the dsDNA of the cell; and / or (B) performing bisulfite conversion of cytosine bases in the plurality of single-stranded DNAs of the plurality of overhanging DNA fragments or the plurality of nucleic acid fragments to generate a plurality of bisulfite-converted ssDNAs containing uracil bases. and / or (ii) determining the methylome information by determining whether the positions of the plurality of barcoded DNA fragments in the sequencing data have a thymine base and whether the corresponding positions in the reference sequence of the dsDNA have a cytosine base, thereby determining that the corresponding positions in the dsDNA have a methylcytosine base; may include The method according to any one of claims 1 to 9.

11. The step of attaching a barcode includes: (a) using the plurality of barcodes to stochastically barcode the plurality of DNA fragments or the plurality of nucleic acid fragments to generate a plurality of stochastically barcoded ssDNA fragments; or (b) barcoding the plurality of DNA fragments or the plurality of nucleic acid fragments using the plurality of barcodes associated with particles to generate the plurality of barcoded ssDNA fragments, wherein the barcodes associated with the particles comprise an identical cellular label sequence and at least 100 different molecular label sequences; (i) at least one barcode of the plurality of barcodes: (A) immobilized on said particles; (B) partially immobilized on said particle; (C) encapsulated in said particles; or (D) Partially encapsulated in the particle It may be (ii) the particles are: (A) is disintegrative; or (B) containing disintegrable hydrogel particles; It may be something, (iii) the particles may comprise a material selected from the group consisting of polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogels, paramagnetic materials, ceramics, plastics, glass, methylstyrene, acrylic polymers, titanium, latex, sepharose, cellulose, nylon, silicone, and any combination thereof; (iv) the barcode of the particle may comprise molecular labels having at least 1,000 different molecular label sequences, and the barcode of the particle may comprise molecular labels having at least 10,000 different molecular label sequences; The method according to any one of claims 1 to 10.

12. the capture sequence comprises a poly(dT) region; and / or the 5' overhang comprises a poly dT sequence; and / or The dsDNA from the cell is selected from the group consisting of nuclear DNA, nucleolar DNA, genomic DNA, mitochondrial DNA, chloroplast DNA, structural DNA, viral DNA, or a combination of two or more of the foregoing listed items. The method according to any one of claims 1 to 11.

13. 1. A kit for sample analysis, comprising: A transposome comprising: a double-stranded nuclease configured to induce double-stranded DNA breaks in a structure containing dsDNA; and Two copies of the adapter with a 5' overhang containing the capture sequence a transposome comprising: a plurality of barcodes, each barcode may include a cell labeling sequence, a molecular labeling sequence, and the capture sequence, at least two of the plurality of barcodes include different molecular labeling sequences, and at least two of the plurality of barcodes include the same cell labeling sequence; Contains kit. (a) the double-stranded nuclease comprises a transposase, a restriction endonuclease, a CRISPR-associated protein, a double-strand-specific nuclease (DSN), or a combination thereof; (b) the transposome further comprises a ligase; (c) the plurality of barcodes comprises at least 10, 50, 100, 500, 1000, 5000, 10,000, 50,000, or 100,000 different molecular labels; and / or (d) said bar code: (i) immobilized on particles, wherein all of the barcodes immobilized on each particle may comprise the same cell-labeling sequence, and different particles may comprise different cell-labeling sequences; or (ii) separated into wells of a substrate, wherein all of the barcodes separated into each well may contain the same cell-labeling sequence, and different wells may contain different cell-labeling sequences; The kit of claim 13.

Citation Information

Patent Citations

  • Transposon end compositions and methods for modifying nucleic acids

    JP2012506704A

  • Massively parallel single-cell analysis

    JP2016533187A

  • Continuity maintained dislocation

    JP2018501776A

  • Single cell whole genome libraries and combinatorial indexing methods of making thereof

    WO2018018008A1

  • Methods of de novo assembly of barcoded genomic DNA fragments

    WO2018031631A1