Single cell chromatin immunoprecipitation sequencing assay

By introducing DNA digestive enzymes and binding reagents into cells and using oligonucleotide barcoding technology, the problem of marking the association status of nuclear targets and DNA at the single cell level is solved, and efficient labeling and measurement of nuclear target-associated DNA is achieved, providing more detailed gene expression and chromatin structure information.

CN120099137APending Publication Date: 2025-06-06BECTON DICKINSON & CO
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510125152.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-07-09
Filing Date
2020-07-21
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to label and measure the association status of nuclear targets with DNA at the single-cell level, limiting high-resolution analysis of gene expression and chromatin structure.

Method used

By permeabilizing the cells and introducing DNA digestive enzymes and specific binding agents into the cells, the double-stranded structure of the nuclear target and DNA is broken down to generate single-stranded overhangs, and these DNA fragments are subsequently barcoded using oligonucleotide barcodes to label and measure DNA associated with the nuclear target.

Benefits of technology

Efficient marking and measurement of nuclear target-associated DNA in single cells is achieved, providing more detailed gene expression and chromatin structure information, and supporting more refined molecular biology research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120099137A_ABST
    Figure CN120099137A_ABST
Patent Text Reader

Abstract

The invention relates to single cell chromatin immunoprecipitation sequencing assays. The disclosure herein includes systems, methods, kits, and compositions for labeling nuclear target associated DNA in a cell. Some embodiments provide digestion compositions comprising a DNA digestive enzyme and a binding agent capable of specifically binding to a nuclear target. Some embodiments provide conjugates comprising a transposon and a binding agent capable of specifically binding to a nuclear target. The transposase may comprise a transposase (e.g., a Tn5 transposase), a first adapter having a first 5 '-protrusion, and a second adapter having a second 5'-protrusion. In some embodiments, a method can include contacting a permeabilized cell comprising a nuclear target associated with dsDNA, such as genomic DNA (gDNA), with a composition provided herein to produce more than one nuclear target associated dsDNA fragment (e.g., a nuclear target associated gDNA fragment), each comprising one or two single-stranded protruding ends.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of an application with a filing date of July 21, 2020, application number 202080048361.7, and invention name “Single Cell Chromatin Immunoprecipitation Sequencing Assay”.

[0002] Related Applications

[0003] This application claims priority under 35 U.S.C. §119(e) to U.S. Provisional Application No. 62 / 876,922, filed on July 22, 2019, and to U.S. Provisional Application No. 63 / 049,980, filed on July 9, 2020. The entire contents of these applications are hereby expressly incorporated herein by reference in their entirety.

[0004] background

[0005] field

[0006] The present disclosure relates generally to the field of molecular biology, such as labeling nuclear targets associated with DNA in single cells.

[0007] Description of the Prior Art

[0008] Current technology allows for measurement of gene expression of single cells in a massively parallel manner (e.g., >10,000 cells) by attaching cell-specific oligonucleotide barcodes to multiple (A) mRNA molecules from individual cells when each cell is co-localized with a barcoding reagent bead in a compartment. There is a need for systems and methods for labeling nuclear target-associated DNA in single cells.

[0009] Overview

[0010] The disclosure herein includes methods of labeling nuclear target associated DNA in a cell. In some embodiments, the method includes: permeabilizing a cell comprising a nuclear target associated with double-stranded deoxyribonucleic acid (dsDNA). The dsDNA can be genomic DNA (gDNA). The method can include: contacting the nuclear target with a digestion composition to produce more than one nuclear target associated dsDNA fragments (e.g., nuclear target associated gDNA fragments) each comprising a single-stranded overhang, the digestion composition comprising a DNA digestion enzyme and a binding agent capable of specifically binding to a nuclear target, wherein each binding agent comprises a binding agent-specific oligonucleotide comprising a unique identifier sequence for the binding agent. The method can include: barcoding more than one nuclear target associated dsDNA fragments or products thereof using a first one or more oligonucleotide barcodes to produce more than one barcoded nuclear target associated DNA fragments, each of the more than one barcoded nuclear target associated DNA fragments comprising a sequence complementary to at least a portion of the nuclear target associated dsDNA fragments, wherein each of the first one or more oligonucleotide barcodes comprises a first target binding region capable of hybridizing to the more than one nuclear target associated dsDNA fragments or products thereof. The barcoded nuclear target associated DNA fragments can be, for example, single-stranded or double-stranded. The method may include: barcoding a binding reagent-specific oligonucleotide or a product thereof using a second one or more oligonucleotide barcodes to produce more than one barcoded binding reagent-specific oligonucleotides, each of the more than one barcoded binding reagent-specific oligonucleotides comprising a sequence complementary to at least a portion of a unique identifier sequence, wherein each of the second one or more oligonucleotide barcodes comprises a second target binding region capable of hybridizing to the binding reagent-specific oligonucleotide or a product thereof.

[0011] In some embodiments, contacting the nuclear target with the digestion composition includes contacting the digestion composition with permeabilized cells. In some embodiments, contacting the nuclear target with the composition includes allowing a binding agent and a DNA digestion enzyme to enter the permeabilized cells and the binding agent to bind to the nuclear target. In some embodiments, the digestion composition comprises a fusion protein containing a binding agent and a digestion enzyme, optionally wherein the size of the fusion protein is set (sized) to diffuse through the nuclear pores of the cell. In some embodiments, the digestion composition comprises a conjugate containing a binding agent and a digestion enzyme, optionally wherein the size of the conjugate is set to diffuse through the nuclear pores of the cell. In some embodiments, the DNA digestion enzyme comprises a domain that can specifically bind to the binding agent. In some embodiments, the domain of the DNA digestion enzyme comprises at least one of protein A, protein G, protein A / G or protein L. In some embodiments, the binding agent and the DNA digestion enzyme are separated from each other when contacting the permeabilized cells, and wherein the binding agent and the DNA digestion enzyme enter the cell separately. In some embodiments, the DNA digestion enzyme binds to the binding agent in the nucleus of the cell. In some embodiments, the DNA digestive enzyme is bound to the binding agent before entering the nucleus of the cell, and the size of the DNA digestive enzyme bound to the binding agent is configured to diffuse through the nuclear pores of the cell. In some embodiments, the DNA digestive enzyme is bound to the binding agent before the binding agent binds to the nuclear target. In some embodiments, the conjugate, fusion protein and / or DNA digestive enzyme bound to the binding agent has a diameter of no more than 120 nm.

[0012] In some embodiments, the first target binding region is complementary to at least a portion of a single-stranded overhang of a nuclear target associated dsDNA fragment. In some embodiments, contacting the nuclear target with a digestion composition produces a complex comprising a binding agent, a DNA digestion enzyme, and nuclear target associated dsDNA fragments, and wherein barcoding more than one nuclear target associated dsDNA fragments comprises contacting the complex with a first more than one oligonucleotide barcode and a second more than one oligonucleotide barcode, optionally comprising digesting the complex with a protease after barcoding, further optionally the protease comprises proteinase K. In some embodiments, the DNA digestion enzyme comprises a restriction enzyme, micrococcal nuclease I (Mnase I), a transposase, a functional fragment thereof, or any combination thereof, optionally the transposase comprises Tn5 transposase, further optionally the digestion composition comprises at least one of a first adaptor having a first 5' overhang and a second adaptor having a second 5' overhang. In some embodiments, barcoding the binding reagent-specific oligonucleotide or its product comprises extending a second one or more oligonucleotide barcode that hybridizes to the binding reagent-specific oligonucleotide or its product, optionally the binding reagent oligonucleotide comprises a sequence complementary to a second target binding region. The method may comprise obtaining sequence data for the one or more barcoded binding reagent-specific oligonucleotides or their products, optionally wherein obtaining sequence information for the one or more barcoded binding reagent-specific oligonucleotides or their products comprises attaching sequencing adapters and / or sequencing primers, their complements and / or portions thereof to the one or more barcoded binding reagent-specific oligonucleotides or their products.

[0013] The disclosure herein includes methods for labeling nuclear target associated DNA in a cell. In some embodiments, the method includes: permeabilizing a cell comprising a nuclear target associated with a double-stranded deoxyribonucleic acid (dsDNA). The dsDNA may be genomic DNA (gDNA). The method may include: contacting the nuclear target with a conjugate to produce more than one nuclear target associated dsDNA fragments each comprising a first 5' overhang and a second 5' overhang, the conjugate comprising a transposome and a binding reagent capable of specifically binding to a nuclear target, wherein the transposome comprises a transposase, a first adaptor having a first 5' overhang, and a second adaptor having a second 5' overhang; and barcoding the more than one nuclear target associated dsDNA fragments or their products using a first more than one oligonucleotide barcode to produce more than one barcoded nuclear target associated DNA fragments, wherein each of the first more than one oligonucleotide barcodes comprises a first target binding region capable of hybridizing with the more than one nuclear target associated dsDNA fragments or their products. Barcoded nuclear target-associated DNA fragments can be, for example, single-stranded or double-stranded.

[0014] In some embodiments, barcoding more than one nuclear target associated dsDNA fragments or products thereof comprises connecting the nuclear target associated dsDNA fragments or products thereof to a first more than one oligonucleotide barcode. In some embodiments, barcoding more than one nuclear target associated dsDNA fragments or products thereof comprises extending the first more than one oligonucleotide barcode hybridized to more than one nuclear target associated dsDNA fragments or products thereof. In some embodiments, the first adapter comprises a first barcode sequence and the second adapter comprises a second barcode sequence. In some embodiments, the first barcode sequence and / or the second barcode sequence identifies a nuclear target. In some embodiments, the first 5' overhang and / or the second 5' overhang comprises a poly(dA) region, a poly(dT) region, or any combination thereof. In some embodiments, the first 5' overhang and / or the second 5' overhang comprises a sequence complementary to the first target binding region or its complement. In some embodiments, the first target binding region comprises a sequence complementary to at least a portion of the 5' region, 3' region, or internal region of the nuclear target associated dsDNA fragments. In some embodiments, the first adaptor and / or the second adaptor comprises a DNA end sequence of a transposon.

[0015] In some embodiments, permeabilization comprises chemical permeabilization or physical permeabilization. In some embodiments, permeabilization comprises contacting the cell with a detergent and / or a surfactant. In some embodiments, permeabilization comprises permeabilizing the cell by sonication. The method may include permeabilizing a nucleus in a cell to produce a permeabilized nucleus. The method may include fixing the cell comprising the nucleus before permeabilizing the nucleus.

[0016] In some embodiments, nuclear targets include methylated nucleotides, DNA-associated proteins, chromatin-associated proteins, or any combination thereof. In some embodiments, the nuclear targets include ALC1, androgen receptor, Bmi-1, BRD4, Brg1, coREST, C-jun, c-Myc, CTCF, EED, EZH2, Fos, histone H1, histone H3, histone H4, heterochromatin protein-1γ, heterochromatin protein-1β, HMGN2 / HMG-17, HP1α, HP1γ, hTERT, Jun, KLF4, K-Ras, Max, MeCP2, MLL / HRX, NPAT, p300, Nanog, NFAT-1, Oct4, P53, Pol II (8WG16), RNA Pol II Ser2P, RNA Pol II Ser5P, RNA Pol II Ser2+5P, RNA Pol II Ser7P, Rb, RNA polymerase II, SMCI, Sox2, STAT1, STAT2, STAT3, Suz12, Tip60, UTF1, or any combination thereof.In some embodiments, the binding agent specifically binds to an epitope comprising a methylated (me), phosphorylated (ph), ubiquitinated (ub), sumoylated (su), biotinylated (bi), or acetylated (ac) histone residue selected from the group consisting of: H1S27ph, H1K25me1, H1K25me2, H1K25me3, H1K26me, H2(A)K4ac, H2(A)K5ac, H2(A)K7ac, H 2(A)S1ph, H2(A)T119ph, H2(A)S122ph, H2(A)S129ph, H2(A)S139ph, H2(A)K119ub, H2(A)K126su, H2(A)K9bi , H2(A)K13bi, H2(B)K5ac, H2(B)K11ac, H2(B)K12ac, H2(B)K15ac, H2(B)K16ac, H2(B)K20ac, H2(B)S10ph, H2( B)S14ph, H2(B)33ph, H2(B)K120ub, H2(B)K123ub, H3K4ac, H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, H3 K56ac, H3K4me1, H3K4me2, H3K4me3, H3R8me, H3K9me1, H3K9me2, H3K9me3, H3R17me, H3K27me1, H3K27me2, H3K 27me3, H3K36me, H3K79me1, H3K79me2, H3K79me3, H3K122ac, H3T3ph, H3S10ph, H3T11ph, H3S28ph, H3K4bi, H3K9bi, H3K18bi, H4K5ac, H4K8ac, H4K12ac, H4K16ac, H4K91ac, H4R3me, H4K20me, H4K59me, H4S1ph, H4K12bi and H4 n-terminal tail ubiquitination or any combination thereof.

[0017] In some embodiments, the binding agent comprises a tetramer, an aptamer, a protein scaffold, or any combination thereof. In some embodiments, the binding agent comprises an antibody or a fragment thereof. In some embodiments, the antibody or a fragment thereof comprises a monoclonal antibody. In some embodiments, the antibody or a fragment thereof comprises a Fab, Fab', F(ab'), 2, Fv, scFv, dsFv, bispecific antibodies (diabody) formed by antibody fragments, triabody, tetraspecific antibody (tetrabody), multispecific antibody (multispecific antibody), single domain antibody (sdAb), single chain comprising complementary scFv (tandem scFv) or bispecific tandem scFv, Fv construct, disulfide bonded Fv, dual variable domain immunoglobulin (DVD-Ig) binding protein or nanobody (nanobody), aptamer, affinity body (affibody), affilin, affitin, affimer, alphabody, anticalin, avimer, DARPin, Fynomer, Kunitz domain peptide, single antibody (monobody) or any combination thereof. In some embodiments, the binding agent is conjugated to the transposome via chemical coupling, genetic fusion, non-covalent association or any combination thereof. In some embodiments, the conjugate is formed by a 1,3-dipolar cycloaddition reaction, a hetero-Diels-Alder reaction, a nucleophilic substitution reaction, a non-aldol carbonyl reaction, a carbon-carbon multiple bond addition, an oxidation reaction, a click reaction, or any combination thereof. In some embodiments, the conjugate is formed by a reaction between acetylene and an azide. In some embodiments, the conjugate is formed by a reaction between an aldehyde or ketone group and a hydrazine or an alkoxyamine. In some embodiments, the binding agent is conjugated to the transposome via at least one of protein A, protein G, protein A / G, or protein L. In some embodiments, the transposase includes a Tn5 transposase. In some embodiments, the transposome comprises a domain that specifically binds to a binding agent, wherein the binding agent and the transposome are separated from each other when in contact with permeabilized cells, and enter the cells separately, and wherein the transposome is bound to the binding agent in the nucleus of the cell. In some embodiments, the transposome is bound to the binding agent before entering the permeabilized cell, and the size of the transposome bound to the binding agent is set to enter the nucleus of the cell through the nuclear pore. In some embodiments, the conjugate has a diameter of no more than 120 nm. In some embodiments, the binding agent and the transposome are each sized to diffuse through the nuclear pores of the cell. In some embodiments, the permeabilized cell comprises an intact nucleus containing chromatin that remains associated with genomic DNA when the binding agent binds to the nuclear target. In some embodiments, binding of the binding agent to the nuclear target occurs in the nucleus of the permeabilized cell.

[0018] The method may include: obtaining sequence data of more than one barcoded nuclear target associated DNA fragment or its product. The method may include: determining information associated with gDNA based on the sequence of more than one barcoded nuclear target associated DNA fragment or its product in the obtained sequencing data. In some embodiments, determining the information associated with gDNA includes determining the genomic information of gDNA based on the sequence of more than one barcoded nuclear target associated DNA fragment in the obtained sequencing data. The method may include: digesting nucleosomes associated with double-stranded gDNA. In some embodiments, determining the genomic information of gDNA includes: determining at least a portion of the sequence of gDNA by aligning the sequence of more than one barcoded nuclear target associated DNA fragment with a reference sequence of gDNA. In some embodiments, determining the information associated with gDNA includes determining the methylome information of gDNA based on the sequence of more than one barcoded nuclear target associated DNA fragment in the obtained sequencing data. The method may include: digesting nucleosomes associated with double-stranded gDNA. In some embodiments, determining the genomic information of gDNA includes: determining at least a portion of the sequence of gDNA by aligning the sequence of more than one barcoded nuclear target associated DNA fragment with a reference sequence of gDNA. In some embodiments, determining the information associated with gDNA includes determining the methylome information of gDNA based on the sequence of more than one barcoded nuclear target associated DNA fragment in the obtained sequencing data. The method may include: digesting nucleosomes associated with double-stranded gDNA. The method may include chemically converting and / or enzymatically converting cytosine bases of more than one nuclear target associated dsDNA fragment or its product to produce more than one converted (e.g., bisulfite converted) nuclear target associated dsDNA fragment with uracil base. In some embodiments, barcoding more than one nuclear target associated dsDNA fragment or its product includes barcoding more than one converted (e.g., bisulfite converted) nuclear target associated dsDNA fragment or its product. The chemical conversion may include bisulfite treatment, and the enzymatic conversion may include APOBEC-mediated conversion. In some embodiments, determining the methylome information includes determining the position of more than one barcoded nuclear target associated DNA fragment with thymine bases in the sequencing data and the corresponding position of the reference sequence of the gDNA with cytosine bases to determine the corresponding position of the gDNA with 5-methylcytosine (5mC) bases and / or 5-hydroxymethylcytosine (5hmC) bases. The method may include: Hi-C, chromatin conformation capture (3C), circularized chromatin conformation capture (4C), carbon copy chromosome conformation capture (5C), chromatin immunoprecipitation (ChIP), ChIP-Loop, combined 3C-ChIP-cloning (6C), capture-C or any combination thereof. The method may include Hi-C / ChiP-seq. The method may include ChiP-seq.

[0019] In some embodiments, the cell contains a copy of a nucleic acid target. The method may include: contacting a second more than one oligonucleotide barcode with a copy of the nucleic acid target for hybridization; extending the second more than one oligonucleotide barcode hybridized with the copy of the nucleic acid target to produce more than one barcoded nucleic acid molecules, each of which contains a sequence complementary to at least a portion of the nucleic acid target and a molecular marker; and obtaining sequence information of more than one barcoded nucleic acid molecules or their products to determine the number of copies of the nucleic acid target in the cell. In some embodiments, extending the first more than one oligonucleotide barcode and / or the second more than one oligonucleotide barcode includes extending the more than one oligonucleotide barcode using a reverse transcriptase and / or a DNA polymerase lacking at least one of a 5' to 3' exonuclease activity and a 3' to 5' exonuclease activity. In some embodiments, the DNA polymerase includes a Klenow fragment. In some embodiments, the reverse transcriptase comprises a viral reverse transcriptase, optionally wherein the viral reverse transcriptase is a murine leukemia virus (MLV) reverse transcriptase or a Moloney murine leukemia virus (MMLV) reverse transcriptase. In some embodiments, the nucleic acid target comprises a nucleic acid molecule. In some embodiments, the nucleic acid molecule comprises ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNA comprising a poly (A) tail, sample index oligonucleotides, cell component binding reagent specific oligonucleotides or any combination thereof. In some embodiments, the first target binding region and / or the second target binding region comprises a poly (dA) region, a poly (dT) region, a random sequence, a gene-specific sequence or any combination thereof.

[0020] In some embodiments, obtaining sequence information of more than one barcoded nuclear target associated DNA fragments includes attaching sequencing adapters and / or sequencing primers, their complementary sequences and / or portions thereof to more than one barcoded nuclear target associated DNA fragments or their products. In some embodiments, obtaining sequence information of more than one barcoded nucleic acid molecules includes attaching sequencing adapters and / or sequencing primers, their complementary sequences and / or portions thereof to more than one barcoded nucleic acid molecules or their products. In some embodiments, each of the first more than one oligonucleotide barcodes comprises a first universal sequence, and each of the second more than one oligonucleotide barcodes comprises a second universal sequence. In some embodiments, the first universal sequence and the second universal sequence are the same. In some embodiments, the first universal sequence and the second universal sequence are different. In some embodiments, the first universal sequence and / or the second universal sequence comprise the following binding sites: sequencing primers and / or sequencing adapters, their complementary sequences and / or portions thereof. In some embodiments, the sequencing adapter includes a P5 sequence, a P7 sequence, a complementary sequence thereof, and / or a portion thereof. In some embodiments, the sequencing primer includes a read 1 sequencing primer, a read 2 sequencing primer, a complementary sequence thereof, and / or a portion thereof. In some embodiments, the first more than one oligonucleotide barcode and the second more than one oligonucleotide barcode each comprise a molecular marker. In some embodiments, at least 10 of the first more than one oligonucleotide barcode and the second more than one oligonucleotide barcode comprise different molecular marker sequences. In some embodiments, each molecular marker of the first more than one oligonucleotide barcode and the second more than one oligonucleotide barcode comprises at least 6 nucleotides. In some embodiments, the first more than one oligonucleotide barcode and the second more than one oligonucleotide barcode are associated with a solid support. In some embodiments, the first more than one oligonucleotide barcode and the second more than one oligonucleotide barcode associated with the same solid support each comprise the same sample tag. In some embodiments, each sample tag of the first more than one oligonucleotide barcode and the second more than one oligonucleotide barcode comprises at least 6 nucleotides. In some embodiments, the first more than one oligonucleotide barcode and the second more than one oligonucleotide barcode each comprise a cell marker. In some embodiments, each cell marker of the first more than one oligonucleotide barcode and the second more than one oligonucleotide barcode comprises at least 6 nucleotides. In some embodiments, the first more than one oligonucleotide barcode and the second more than one oligonucleotide barcode associated with the same solid support comprise the same cell marker. In some embodiments, the first more than one oligonucleotide barcode and the second more than one oligonucleotide barcode associated with different solid supports comprise different cell markers.

[0021] In some embodiments, the solid support comprises synthetic particles, a flat surface, or a combination thereof. The method may include: associating synthetic particles comprising a first more than one oligonucleotide barcode and a second more than one oligonucleotide barcode with a cell. The method may include: lysing the cell after associating the synthetic particle with the cell, optionally lysing the cell comprises heating the cell, contacting the cell with a detergent, changing the pH of the cell, or any combination thereof. In some embodiments, the synthetic particle and the single cell are in the same partition, and optionally the partition is a well or a droplet. In some embodiments, at least one of the first more than one oligonucleotide barcode and the second more than one oligonucleotide barcode is fixed or partially fixed on the synthetic particle, or at least one of the first more than one oligonucleotide barcode and the second more than one oligonucleotide barcode is encapsulated or partially encapsulated in the synthetic particle. In some embodiments, the synthetic particle is destructible, optionally a destructible hydrogel particle. In some embodiments, the synthetic particles include beads, optionally beads include sepharose beads, streptavidin beads, agarose beads, magnetic beads, conjugated beads, protein A conjugated beads, protein G conjugated beads, protein A / G conjugated beads, protein L conjugated beads, oligo(dT) conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, or any combination thereof. In some embodiments, the synthetic particles include a material selected from the group consisting of polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic substances, ceramics, plastics, glass, methylstyrene, acrylic polymers, titanium, latex, agarose gel, cellulose, nylon, silicone, and any combination thereof. In some embodiments, each of the first more than one oligonucleotide barcode and the second more than one oligonucleotide barcode comprises a linker functional group. In some embodiments, the synthetic particles include a solid support functional group. In some embodiments, the support functional group and the linker functional group are associated with each other, and optionally the linker functional group and the support functional group are individually selected from the group consisting of C6, biotin, streptavidin, one or more primary amines, one or more aldehydes, one or more ketones, and any combination thereof.

[0022] The disclosure herein includes kits. The kits may include a digestion composition comprising a DNA digestion enzyme and a binding agent capable of specifically binding to a nuclear target, wherein the binding agent comprises a binding agent-specific oligonucleotide comprising a unique identifier sequence for the binding agent.

[0023] In some embodiments, the DNA digestion enzyme comprises a restriction enzyme, micrococcal nuclease I, a transposase, a functional fragment thereof, or any combination thereof, optionally, the transposase comprises a Tn5 transposase, and further optionally the digestion composition comprises at least one of a first adaptor having a first 5' overhang and a second adaptor having a second 5' overhang. In some embodiments, the digestion composition comprises a fusion protein comprising a binding agent and a digestion enzyme, optionally wherein the size of the fusion protein is configured to diffuse through the nuclear pores of the cell. In some embodiments, the digestion composition comprises a conjugate comprising a binding agent and a digestion enzyme, optionally wherein the size of the conjugate is configured to diffuse through the nuclear pores of the cell. In some embodiments, the DNA digestion enzyme comprises a domain capable of specifically binding to a binding agent, optionally the domain of the DNA digestion enzyme comprises at least one of protein A, protein G, protein A / G, or protein L, and further optionally the size of the DNA digestion enzyme bound to the binding agent is configured to diffuse through the nuclear pores of the cell. The kit may comprise: a protease (e.g., proteinase K).

[0024] The disclosure herein includes a kit. In some embodiments, the kit comprises: a conjugate comprising a transposome and a binding agent capable of specifically binding to a nuclear target, wherein the transposome comprises a transposase, a first adapter having a first 5' overhang, and a second adapter having a second 5' overhang. The kit may comprise: a DNA polymerase lacking at least one of a 5' to 3' exonuclease activity and a 3' to 5' exonuclease activity, optionally wherein the DNA polymerase comprises a Klenow fragment. The kit may comprise: a reverse transcriptase, optionally wherein the reverse transcriptase comprises a viral reverse transcriptase, and optionally the viral reverse transcriptase is a murine leukemia virus (MLV) reverse transcriptase or a Moloney murine leukemia virus (MMLV) reverse transcriptase. The kit may comprise: a ligase. The kit may comprise: a detergent and / or a surfactant. The kit may comprise: a buffer, a cartridge, or both. The kit may comprise: one or more reagents for reverse transcription reactions and / or amplification reactions.

[0025] In some embodiments, the nuclear target comprises a DNA-associated protein or a chromatin-associated protein. In some embodiments, the nuclear target comprises a methylated nucleotide. In some embodiments, the nuclear targets include ALC1, androgen receptor, Bmi-1, BRD4, Brg1, coREST, C-jun, c-Myc, CTCF, EED, EZH2, Fos, histone H1, histone H3, histone H4, heterochromatin protein-1γ, heterochromatin protein-1β, HMGN2 / HMG-17, HP1α, HP1γ, hTERT, Jun, KLF4, K-Ras, Max, MeCP2, MLL / HRX, NPAT, p300, Nanog, NFAT-1, Oct4, P53, Pol II (8WG16), RNA Pol II Ser2P, RNA Pol II Ser5P, RNA Pol II Ser2+5P, RNA Pol II Ser7P, Rb, RNA polymerase II, SMCI, Sox2, STAT1, STAT2, STAT3, Suz12, Tip60, UTF1, or any combination thereof.In some embodiments, the binding agent specifically binds to an epitope comprising a methylated (me), phosphorylated (ph), ubiquitinated (ub), paraubiquitinated (su), biotinylated (bi) or acetylated (ac) histone residue selected from the group consisting of: H1S27ph, H1K25me1, H1K25me2, H1K25me3, H1K26me, H2(A)K4ac, H2(A)K5ac, H2(A)K7ac, H2(A)S1ph, H2(A)T119ph, H2(A)S122ph, H2(A)S129ph, H2(A)S139ph, H2(A)K119ub, H2(A)K126su, H2(A)K9bi, H2(A)K1 3bi, H2(B)K5ac, H2(B)K11ac, H2(B)K12ac, H2(B)K15ac, H2(B)K16ac, H2(B)K20ac, H2(B)S10ph, H2(B)S14p h, H2(B)33ph, H2(B)K120ub, H2(B)K123ub, H3K4ac, H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, H3K56a c. H3K4me1, H3K4me2, H3K4me3, H3R8me, H3K9me1, H3K9me2, H3K9me3, H3R17me, H3K27me1, H3K27me2, H3K27m e3, H3K36me, H3K79me1, H3K79me2, H3K79me3, H3K122ac, H3T3ph, H3S10ph, H3T11ph, H3S28ph, H3K4bi, H3K9bi, H3K18bi, H4K5ac, H4K8ac, H4K12ac, H4K16ac, H4K91ac, H4R3me, H4K20me, H4K59me, H4S1ph, H4K12bi and H4 n-terminal tail ubiquitination or any combination thereof. In some embodiments, the binding agent comprises a tetramer, an aptamer, a protein scaffold or any combination thereof. In some embodiments, the binding agent comprises an antibody or a fragment thereof. In some embodiments, the antibody or a fragment thereof comprises a monoclonal antibody. In some embodiments, wherein the antibody or a fragment thereof comprises Fab, Fab', F(ab'). 2, Fv, scFv, dsFv, bispecific antibodies formed by antibody fragments, trispecific antibodies, tetraspecific antibodies, multispecific antibodies, single domain antibodies (sdAb), single chains comprising complementary scFv (tandem scFv) or bispecific tandem scFv, Fv constructs, disulfide-linked Fv, dual variable domain immunoglobulin (DVD-Ig) binding proteins or nanobodies, aptamers, affinity bodies, affilins, affitins, affimers, alphabodies, anticalins, avimers, DARPins, Fynomers, Kunitz domain peptides, single antibodies or any combination thereof. In some embodiments, the binding agent is conjugated to the transposome via chemical coupling, genetic fusion, non-covalent association or any combination thereof.

[0026] In some embodiments, the conjugate is formed by a 1,3-dipolar cycloaddition reaction, a hetero-Diels-Alder reaction, a nucleophilic substitution reaction, a non-aldol carbonyl reaction, a carbon-carbon multiple bond addition, an oxidation reaction, a click reaction, or any combination thereof. In some embodiments, the conjugate is formed by a reaction between acetylene and an azide. In some embodiments, the conjugate is formed by a reaction between an aldehyde or ketone group and a hydrazine or an alkoxyamine. In some embodiments, the binding agent is conjugated to the transposome via at least one of protein A, protein G, protein A / G, or protein L. In some embodiments, the transposase includes a Tn5 transposase. In some embodiments, the conjugate has a diameter of no more than 120 nm. In some embodiments, the size of each of the binding agent and the transposome is configured to diffuse through the nuclear pore.

[0027] The kit may include: more than one oligonucleotide barcode. In some embodiments, each of the more than one oligonucleotide barcode comprises a molecular marker and a target binding region. In some embodiments, at least 10 of the more than one oligonucleotide barcodes comprise different molecular marker sequences. In some embodiments, the target binding region comprises a gene-specific sequence, an oligo(dT) sequence, a random polymer, or any combination thereof. The oligonucleotide barcodes in the first more than one oligonucleotide barcode and / or the second more than one oligonucleotide barcode may comprise a target binding region (e.g., a first target binding region) that can hybridize (e.g., be complementary) with a dsDNA fragment or its product associated with more than one nuclear target, such as a 5' overhang and / or a 3' overhang (e.g., a 5' overhang of the first adapter, a 5' overhang of the second adapter). In some embodiments, the oligonucleotide barcodes comprise the same sample marker and / or the same cell marker. In some embodiments, each sample marker and / or cell marker of more than one oligonucleotide barcode comprises at least 6 nucleotides. In some embodiments, each molecular marker of more than one oligonucleotide barcode comprises at least 6 nucleotides.

[0028] In some embodiments, at least one of the more than one oligonucleotide barcodes is partially fixed on a synthetic particle. In some embodiments, at least one of the more than one oligonucleotide barcodes is fixed or partially fixed on a synthetic particle; and / or at least one of the more than one oligonucleotide barcodes is encapsulated or partially encapsulated in a synthetic particle. In some embodiments, the synthetic particles include beads. In some embodiments, the beads include agarose gel beads, streptavidin beads, agarose beads, magnetic beads, conjugated beads, protein A conjugated beads, protein G conjugated beads, protein A / G conjugated beads, protein L conjugated beads, oligo(dT) conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, or any combination thereof. In some embodiments, the synthetic particles comprise a material selected from the group consisting of polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic substances, ceramics, plastics, glass, methylstyrene, acrylic polymers, titanium, latex, agarose gel, cellulose, nylon, silicone, and any combination thereof. In some embodiments, each of more than one oligonucleotide barcode comprises a linker functional group. In some embodiments, the synthetic particles comprise a solid support functional group. In some embodiments, the support functional group and the linker functional group are associated with each other. In some embodiments, the linker functional group and the support functional group are individually selected from the group consisting of C6, biotin, streptavidin beads, one or more primary amines, one or more aldehydes, one or more ketones, and any combination thereof.

[0029] In some embodiments, a kit is described. The kit may comprise a binding reagent (e.g., a protein binding reagent) associated with a reagent oligonucleotide comprising a unique identifier for the protein binding reagent, wherein the protein binding reagent specifically binds to a target protein. The kit may comprise a fusion protein comprising a DNA digesting enzyme, wherein (i) the fusion protein comprises a domain that specifically binds to the protein binding reagent, or (ii) the fusion protein binds to the protein binding reagent. The kit may comprise a first oligonucleotide probe comprising a first target binding region and a sample identifier sequence, wherein the first target binding region and at least a portion of the reagent oligonucleotide associated with the protein binding reagent are complementary. The kit may comprise a second oligonucleotide probe comprising a second target binding region and a sample identifier sequence, wherein the second target binding region is complementary to at least a portion of the DNA. In some embodiments, for any of the kits described herein, the target protein comprises a DNA-associated protein or a chromatin-associated protein. In some embodiments, for any of the kits described herein,The target protein is selected from the group consisting of ALC1, androgen receptor, Bmi-1, BRD4, Brg1, coREST, C-jun, c-Myc, CTCF, EED, EZH2, Fos, histone H1, histone H3, histone H4, heterochromatin-1γ, heterochromatin-1β, HMGN2 / HMG-17, HP1α, HP1γ, hTERT, Jun, KLF4, K-Ras, Max, MeCP2, MLL / HRX, NPAT, p300, Nanog, NFAT-1, Oct4, P53, Pol II (8WG16), RNA Pol II Ser2P, RNA Pol II Ser5P, RNA Pol II Ser2+5P, RNA Pol II Ser7P, Rb, RNA polymerase II, SMCI, Sox2, STAT1, STAT2, STAT3, Suz12, Tip60, UTF1, H1S27ph, H1K25me1, H1K25me2, H1K25me3, H1K26me, H2(A)K4ac, H2(A)K5ac, H2(A)K7ac, H2(A)S1ph, H2(A)T119ph, H2(A)S122ph , H2(A)S129ph, H2(A)S139ph, H2(A)K119ub, H2(A)K126su, H2(A)K9bi, H2(A)K13bi, H2(B)K5ac, H2(B) K11ac, H2(B)K12ac, H2(B)K15ac, H2(B)K16ac, H2(B)K20ac, H2(B)S10ph, H2(B)S14ph, H2(B)33ph, H2( B)K120ub, H2(B)K123ub, H3K4ac, H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, H3K56ac, H3K4me1, H3 K4me2, H3K4me3, H3R8me, H3K9me1, H3K9me2, H3K9me3, H3R17me, H3K27me1, H3K27me2, H3K27me3, H3K36 me, H3K79me1, H3K79me2, H3K79me3, H3K122ac, H3T3ph, H3S10ph, H3T11ph, H3S28ph, H3K4bi, H3K9bi, H3K18bi, H4K5ac, H4K8ac, H4K12ac, H4K16ac, H4K91ac, H4R3me, H4K20me, H4K59me, H4S1ph, H4K12bi, and H4 N-terminal tail ubiquitination,Or a combination of two or more of the listed items. In some embodiments, for any kit described herein, the first oligonucleotide probe and the second oligonucleotide probe comprise the same sample identifier sequence. In some embodiments, for any kit described herein, the first oligonucleotide probe and the second oligonucleotide probe are fixed to a substrate. In some embodiments, for any kit described herein, the substrate is a solid surface, a bead, a microarray, a plate, a tube, or a well. In some embodiments, for any kit described herein, the kit further comprises a composition containing reagents for generating a nucleic acid library, wherein the reagents comprise at least one of a DNA restriction enzyme, a nuclease (such as micrococcal nuclease), a lysozyme, a proteinase K, a random hexamer, a polymerase (Φ29 DNA polymerase, Taq polymerase, Bsu polymerase, Klenow polymerase), a transposase (Tn5), a primer (P5 and P7 adapter sequences), a ligase, a catalytic enzyme, a deoxynucleotide triphosphate, a buffer, or a divalent cation. ,

[0030] In some embodiments, a method for labeling the target protein associated DNA of a single cell is described. The method may include permeabilizing a cell comprising a target protein. The method may include contacting the permeabilized cell with a digestion composition, the digestion composition comprising a binding agent (e.g., a protein binding agent) associated with a reagent oligonucleotide, the reagent oligonucleotide comprising a unique identifier for a protein binding agent, wherein the protein binding agent specifically binds to the target protein. The digestion composition may include a DNA digestion enzyme and a binding agent that can specifically bind to a nuclear target (e.g., a nucleoprotein). The composition may also include a fusion protein containing a DNA digestion enzyme, wherein (i) the fusion protein also includes a domain that specifically binds to a protein binding agent, or (ii) the fusion protein binds to a protein binding agent. The protein binding agent and the fusion protein may enter the permeabilized cell. The method may include contacting the target protein of the permeabilized cell with a protein binding agent and a fusion protein bound to the protein binding agent. The method may include binding a binding agent (e.g., a protein binding agent) to a target protein in a cell, wherein the DNA of the cell associates with the target protein. The method may include digesting the DNA associated with the target protein with a DNA digesting enzyme of the fusion protein, the digested DNA comprising a single-stranded overhang, thereby forming a complex, the complex comprising: a protein binding reagent; a fusion protein; and the digested DNA. The method may include contacting the complex with a first oligonucleotide probe and a second oligonucleotide probe, the first oligonucleotide probe comprising a first target binding region and a sample identifier sequence, wherein the first target binding region is complementary to at least a portion of the reagent oligonucleotide; the second oligonucleotide probe comprises a second target binding region and a sample identifier sequence. The second target binding region may be complementary to at least a portion of the single-stranded overhang of the DNA. The first oligonucleotide probe and the second oligonucleotide probe may comprise the same sample identifier sequence. The method may include extending the oligonucleotide probe to produce more than one labeled nucleic acid, the labeled nucleic acid comprising a reverse complement of the sample identifier sequence and at least a portion of the reagent oligonucleotide or at least a portion of the DNA. In some embodiments, for any method described herein, (i) the fusion protein comprises a domain that specifically binds to a protein binding reagent, wherein the protein binding reagent and the fusion protein separate from each other when contacting the permeabilized cell and enter the cell separately, and wherein the fusion protein binds to the protein binding reagent in the nucleus of the cell. In some embodiments, for any of the methods described herein, the domain of the fusion protein comprises at least one of protein A, protein G, protein A / G, or protein L. In some embodiments, for any of the methods described herein, (ii) the fusion protein is bound to a protein binding agent, and the size of the fusion protein bound to the protein binding agent is configured to enter the nucleus of the cell through a nuclear pore. In some embodiments, for any of the methods described herein, the method further comprises a multiplex construction comprising the use of two or more different protein binding agents, each of which is specific for a different target protein.In some embodiments, for any method described herein, the fusion protein bound to the protein binding reagent has a diameter of no more than 120nm. In some embodiments, for any method described herein, the size of each of the protein binding reagent and the fusion protein is set to diffuse through the nuclear pores of the cell. In some embodiments, for any method described herein, the first oligonucleotide probe and the second oligonucleotide probe are fixed on a substrate. In some embodiments, for any method described herein, the substrate includes at least one of a bead, a microarray, a plate, a tube or a hole. In some embodiments, for any method described herein, the method also includes capturing the complex on the substrate. In some embodiments, for any method described herein, the first oligonucleotide probe includes a first barcode sequence, and the second oligonucleotide probe includes a second barcode sequence. The first barcode sequence and the second barcode sequence can each be from a set of different unique barcode sequences. In some embodiments, for any method described herein, the permeabilized cell includes a complete nucleus, and the nucleus includes a chromatin that is associated with genomic DNA when the protein binding reagent is bound to the target protein. In some embodiments, for any method described herein, the protein binding reagent includes an antibody that specifically binds to the target protein. In some embodiments, for any method described herein, binding of a protein binding agent to a target protein and digestion of DNA associated with the target protein occurs in the nucleus of a cell. In some embodiments, for any method described herein, the target protein is selected from the group consisting of ALC1, androgen receptor, Bmi-1, BRD4, Brg1, coREST, C-jun, c-Myc, CTCF, EED, EZH2, Fos, histone H1, histone H3, histone H4, heterochromatin protein-1γ, heterochromatin protein-1β, HMGN2 / HMG-17, HP1α, HP1γ, hTERT, Jun, KLF4, K-Ras, Max, MeCP2, MLL / HRX, NPAT, p300, Nanog, NFAT-1, Oct4, P53, Pol II (8WG16), RNA Pol II Ser2P, RNA Pol II Ser5P, RNA Pol II Ser2+5P, RNA Pol II Ser7P, Rb, RNA polymerase II, SMCI, Sox2, STAT1, STAT2, STAT3, Suz12, Tip60 and UTF1, or a combination of two or more of the listed items.In some embodiments, for any of the methods described herein, the protein binding agent specifically binds to an epitope that comprises, consists essentially of, or consists of a phosphorylated (ph), methylated (me), panthenylated (ub), threonylated (su), biotinylated (bi), or acetylated (ac) histone residue selected from the group consisting of: H1S27ph, H1K25me1, H1K25me2, H1K25me3, H1K26me, H2(A )K4ac, H2(A)K5ac, H2(A)K7ac, H2(A)S1ph, H2(A)T119ph, H2(A)S122ph, H2(A)S129ph, H2(A)S139ph, H2(A)K119ub , H2(A)K126su, H2(A)K9bi, H2(A)K13bi, H2(B)K5ac, H2(B)K11ac, H2(B)K12ac, H2(B)K15ac, H2(B)K16ac, H2(B)K2 0ac, H2(B)S10ph, H2(B)S14ph, H2(B)33ph, H2(B)K120ub, H2(B)K123ub, H3K4ac, H3K9ac, H3K14ac, H3K18ac, H3K2 3ac, H3K27ac, H3K56ac, H3K4me1, H3K4me2, H3K4me3, H3R8me, H3K9me1, H3K9me2, H3K9me3, H3R17me, H3K27me1, H3K 27me2, H3K27me3, H3K36me, H3K79me1, H3K79me2, H3K79me3, H3K122ac, H3T3ph, H3S10ph, H3T11ph, H3S28ph, H3K4bi, H3K9bi, H3K18bi, H4K5ac, H4K8ac, H4K12ac, H4K16ac, H4K91ac, H4R3me, H4K20me, H4K59me, H4S1ph, H4K12bi and H4 n-terminal tail ubiquitination, or a combination of two or more of the listed items. In some embodiments, for any method described herein, permeabilization comprises chemical permeabilization or physical permeabilization. In some embodiments, for any of the methods described herein, the method further comprises permeabilizing the cells by contacting the cells with a detergent or surfactant (such as guanidine hydrochloride, Triton X, digitonin, or a combination thereof). In some embodiments, for any of the methods described herein, the method further comprises permeabilizing the cells by sonication. In some embodiments, for any of the methods described herein, the DNA digesting enzyme comprises a restriction enzyme, micrococcal nuclease I, or a transposase (e.g., Tn5 transposase) or a functional fragment thereof.The transposase can be a Tn transposase (e.g., Tn3, Tn5, Tn7, Tn10, Tn552, Tn903), a MuA transposase, a Vibhar transposase (e.g., from Vibrio harveyi), Ac-Ds, Ascot-1, Bs1, Cin4, Copia, En / Spm, an F element, hobo, Hsmar1, Hsmar2, IN (HIV), IS1, IS2, IS3, IS4, IS5, IS6, IS10, IS21, IS30, IS50, IS51, IS150, IS256, IS407, IS427, IS630, IS903, IS911, IS982, IS1031, ISL2, L1, Mariner, P element, Tam3, Tc1, Tc3, Tel, THE-1, Tn / O, TnA, Tn3, Tn5, Tn7, Tn10, Tn552, Tn903, Tol1, Tol2, Tn10, Ty1, any prokaryotic transposase, or any transposase related to or derived from those listed above. In some embodiments, the transposase related to and / or derived from a parent transposase can comprise a peptide fragment having at least about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% amino acid sequence homology with the corresponding peptide fragment of the parent transposase. The length of the peptide fragment can be at least about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 60, about 70, about 80, about 90, about 100, about 150, about 200, about 250, about 300, about 400, or about 500 amino acids. For example, a transposase derived from Tn5 can comprise a peptide fragment that is 50 amino acids in length and is about 80% homologous to the corresponding fragment in the parent Tn5 transposase. In some cases, insertion can be promoted and / or triggered by adding one or more cations. The cation can be a divalent cation, such as Ca. 2+ Mg 2+ and Mn 2+. In some embodiments, for any of the methods described herein, digestion of DNA is performed in the nucleus of a permeabilized cell. In some embodiments, for any of the methods described herein, the first target binding region of the first oligonucleotide probe comprises a poly-T sequence, and the reagent oligonucleotide comprises a poly-T sequence. In some embodiments, for any of the methods described herein, the first target binding region of the first oligonucleotide probe does not comprise five consecutive thymidines, and the reagent oligonucleotide comprises a sequence complementary to the first target binding region of the first oligonucleotide probe. In some embodiments, for any of the methods described herein, the second target binding region of the second oligonucleotide probe comprises a sequence complementary to at least a portion of the 5' region, 3' region, or internal region of the DNA. In some embodiments, for any of the methods described herein, the method further comprises connecting the second target binding region to the digested DNA. In some embodiments, for any of the methods described herein, the method further comprises, for example, if the reagent oligonucleotide comprises single-stranded DNA or RNA, using a reverse transcriptase to produce a target-barcode conjugate. In some embodiments, for any of the methods described herein, the method further comprises producing a nucleic acid library comprising more than one labeled nucleic acid. In some embodiments, for any of the methods described herein, producing a nucleic acid library comprises contacting with a reagent comprising an enzyme, a chemical substance, or a primer. In some embodiments, for any of the methods described herein, the reagents include at least one of: a DNA restriction enzyme, a nuclease (such as micrococcal nuclease), a lysozyme, a proteinase K, a random hexamer, a polymerase (Φ29 DNA polymerase, Taq polymerase, Bsu polymerase, Klenow polymerase), a transposase (Tn5), a primer (P5 and P7 adapter sequences), a ligase, a catalytic enzyme, a deoxynucleotide triphosphate, a buffer, or a divalent cation. In some embodiments, for any of the methods described herein, the method further includes digesting the protein binding reagent, the fusion protein, and the target protein with a protease after the extension. In some embodiments, for any of the methods described herein, the protease includes a proteinase K. In some embodiments, for any of the methods described herein, the method further includes connecting a double-stranded DNA (dsDNA) template to a second oligonucleotide probe. The dsDNA template can include a template strand and a complementary strand. In some embodiments, for any of the methods described herein, the template strand includes a template switching oligonucleotide and a unique capture sequence, and the complementary strand includes a sequence complementary to the unique capture sequence. In some embodiments, for any method described herein, the unique capture sequence comprises a nucleic acid sequence of no more than 40 nucleotides in length. In some embodiments, for any method described herein, the dsDNA template further comprises a randomer. In some embodiments, for any method described herein, the method further comprises denaturing the dsDNA template and removing the template strand, wherein the complementary strand remains connected to the second oligonucleotide probe.In some embodiments, for any of the methods described herein, the dsDNA template is attached to biotin, and removing the template strand comprises contacting with streptavidin.

[0031] In some embodiments, a method for attaching a target binding region to an oligonucleotide probe is described. The method may include providing an oligonucleotide probe, wherein the oligonucleotide probe includes a sample identifier sequence and a hybridization domain. The method may include contacting the oligonucleotide probe with a partially double-stranded DNA (dsDNA) template comprising a template strand and a complementary strand. The template strand may include a single-stranded complementary hybridization domain complementary to at least a portion of the hybridization domain, and a sequence complementary to the target binding region. The complementary strand may include a complementary strand containing a target binding region, wherein the target binding region of the complementary strand hybridizes to a sequence complementary to the target binding region of the template strand. The method may include connecting the dsDNA template to the oligonucleotide probe. The method may include denaturing the dsDNA template. The method may include removing the template strand. In some embodiments, for any of the methods described herein, the oligonucleotide probe is associated with a substrate. In some embodiments, for any of the methods described herein, the substrate includes at least one of a bead, a microarray, a plate, a tube, or a well. In some embodiments, for any of the methods described herein, the template strand is associated with biotin, and wherein removing the template strand includes contacting biotin with streptavidin. In some embodiments, for any of the methods described herein, the hybridization domain of the oligonucleotide probe comprises a poly-T sequence, and wherein the complementary hybridization domain comprises a poly-A sequence. In some embodiments, for any of the methods described herein, the target binding region does not comprise a poly-T sequence. In some embodiments, for any of the methods described herein, the method further comprises appending a first target binding region to the first oligonucleotide probe and / or appending a second target binding region to the second oligonucleotide probe according to the methods described herein.

[0032] In some embodiments, a test kit is described. The test kit may include an oligonucleotide probe containing a sample identifier sequence and a hybridization domain. The test kit may include a DNA template, and the DNA template includes a template strand and a complementary strand. The template strand may include a single-stranded complementary hybridization domain (complementary to at least a portion of the hybridization domain) and a sequence complementary to the target binding region. The complementary strand may include a target binding region. Optionally, in the test kit, the target binding region is hybridized with a sequence complementary to the target binding region, so that the template strand is hybridized with the complementary strand. Therefore, the template strand and the complementary strand can form a partially double-stranded DNA with at least a portion of the hybridization domain of the single strand. In some embodiments, for any test kit described herein, the template strand sequence is associated with biotin, and wherein the test kit further includes a composition containing streptavidin. In some embodiments, for any test kit described herein, the oligonucleotide probe is associated with a substrate. In some embodiments, for any test kit described herein, the substrate includes at least one of a bead, a microarray, a plate, a tube, or a well. In some embodiments, for any test kit described herein, the test kit further includes a denaturant. In some embodiments, for any kit described herein, the denaturing agent comprises dimethyl sulfoxide (DMSO), formamide, guanidine, propylene glycol, salt, sodium hydroxide, sodium salicylate or urea. In some embodiments, any kit described in this paragraph further comprises any kit described herein. In some embodiments, for any kit or method or composition described herein, the protein binding reagent comprises an antibody.

[0033] In some embodiments, compositions are described. The composition may include cells containing a target protein associated with DNA. The composition may include a complex bound to the target protein. The complex may include a protein binding reagent associated with a reagent oligonucleotide, and the reagent oligonucleotide includes a unique identifier for a protein binding reagent. The protein binding reagent can specifically bind to the protein of interest. The complex may include a fusion protein containing a DNA digestion enzyme bound to the protein binding reagent. In some embodiments, for any composition described herein, the fusion protein includes a domain that specifically binds to the protein binding reagent. In some embodiments, for any composition described herein, the complex is bound to the target protein in the nucleus of the cell. In some embodiments, for any composition described herein, the DNA digestion enzyme includes a restriction enzyme, micrococcal nuclease I or Tn5 transposase or a functional fragment thereof.

[0034] The present invention also provides the following embodiments:

[0035] 1. A method for labeling nuclear target-associated DNA in a cell, the method comprising:

[0036] permeabilizing a cell comprising a nuclear target associated with double-stranded deoxyribonucleic acid (dsDNA), wherein optionally the dsDNA is genomic DNA (gDNA);

[0037] contacting the nuclear target with a digestion composition comprising a DNA digestion enzyme and a binding agent capable of specifically binding to the nuclear target, wherein each binding agent comprises a binding agent-specific oligonucleotide comprising a unique identifier sequence for the binding agent, to produce more than one nuclear target-associated dsDNA fragments each comprising a single-stranded overhang;

[0038] barcoding the more than one nuclear target associated dsDNA fragments or products thereof using first one or more oligonucleotide barcodes to generate more than one barcoded nuclear target associated DNA fragments, each of the more than one barcoded nuclear target associated DNA fragments comprising a sequence complementary to at least a portion of the nuclear target associated dsDNA fragments, wherein each of the first one or more oligonucleotide barcodes comprises a first target binding region capable of hybridizing to the more than one nuclear target associated dsDNA fragments or products thereof; and

[0039] The binding reagent-specific oligonucleotide or its product is barcoded using a second one or more oligonucleotide barcode to produce more than one barcoded binding reagent-specific oligonucleotide, each of the more than one barcoded binding reagent-specific oligonucleotides comprising a sequence complementary to at least a portion of the unique identifier sequence, wherein each of the second one or more oligonucleotide barcodes comprises a second target binding region capable of hybridizing to the binding reagent-specific oligonucleotide or its product.

[0040] 2. The method of embodiment 1, wherein contacting the nuclear target with the digestion composition comprises contacting the digestion composition with permeabilized cells.

[0041] 3. The method of any one of embodiments 1-2, wherein contacting the nuclear target with the composition comprises entering the binding agent and the DNA digesting enzyme into the permeabilized cell and the binding agent binding to the nuclear target.

[0042] 4. The method of any one of embodiments 1-3, wherein the digestion composition comprises a fusion protein comprising the binding agent and the digestive enzyme, optionally wherein the fusion protein is sized to diffuse through the nuclear pores of the cell.

[0043] 5. The method of any one of embodiments 1-3, wherein the digestive composition comprises a conjugate comprising the binding agent and the digestive enzyme, optionally wherein the conjugate is sized to diffuse through the nuclear pores of the cell.

[0044] 6. A method according to any one of embodiments 1-3, wherein the DNA digesting enzyme comprises a domain capable of specifically binding to the binding agent, and optionally the domain of the DNA digesting enzyme comprises at least one of protein A, protein G, protein A / G or protein L.

[0045] 7. A method according to embodiment 6, wherein the binding reagent and the DNA digesting enzyme are separated from each other when contacting the permeabilized cells, and wherein the binding reagent and the DNA digesting enzyme enter the cells separately.

[0046] 8. The method of any one of embodiments 6-7, wherein the DNA digesting enzyme binds to the binding agent within the nucleus of the cell.

[0047] 9. A method according to any one of embodiments 6-8, wherein the DNA digesting enzyme binds to the binding reagent before entering the nucleus of the cell, and wherein the size of the DNA digesting enzyme bound to the binding reagent is configured to diffuse through the nuclear pores of the cell.

[0048] 10. The method of any one of embodiments 6-9, wherein the DNA digesting enzyme binds to the binding agent before the binding agent binds to the nuclear target.

[0049] 11. The method according to any one of embodiments 1-10, wherein the conjugate, the fusion protein and / or the DNA digesting enzyme bound to the binding agent has a diameter of no more than 120 nm.

[0050] 12. The method of any one of embodiments 1-11, wherein the first target binding region is complementary to at least a portion of the single-stranded overhang of the nuclear target-associated dsDNA fragment.

[0051] 13. A method according to any one of embodiments 1-12, wherein contacting the nuclear target with the digestion composition produces a complex comprising the binding reagent, the DNA digestion enzyme, and nuclear target-associated dsDNA fragments, and wherein barcoding the more than one nuclear target-associated dsDNA fragments comprises contacting the complex with the first more than one oligonucleotide barcode and the second more than one oligonucleotide barcode, optionally comprising digesting the complex with a protease after barcoding, and further optionally the protease comprises proteinase K.

[0052] 14. The method of any one of embodiments 1-13, wherein the DNA digestion enzyme comprises a restriction enzyme, micrococcal nuclease I, a transposase, a functional fragment thereof, or any combination thereof, optionally the transposase comprises Tn5 transposase, further optionally the digestion composition comprises at least one of a first adaptor having a first 5' overhang and a second adaptor having a second 5' overhang.

[0053] 15. A method according to any one of embodiments 1-14, wherein barcoding the binding reagent-specific oligonucleotide or its product includes extending the second one or more oligonucleotide barcodes hybridized to the binding reagent-specific oligonucleotide or its product, optionally wherein the binding reagent oligonucleotide comprises a sequence complementary to the second target binding region.

[0054] 16. A method according to any one of embodiments 1-15, the method comprising obtaining sequence data of the more than one barcoded binding reagent-specific oligonucleotides or their products, optionally wherein obtaining sequence information of the more than one barcoded binding reagent-specific oligonucleotides or their products comprises attaching sequencing adapters and / or sequencing primers, their complementary sequences and / or portions thereof to the more than one barcoded binding reagent-specific oligonucleotides or their products.

[0055] 17. A method for labeling nuclear target-associated DNA in a cell, the method comprising:

[0056] permeabilizing a cell comprising a nuclear target associated with double-stranded deoxyribonucleic acid (dsDNA), wherein optionally the dsDNA is genomic DNA (gDNA);

[0057] contacting the nuclear target with a conjugate comprising a transposome and a binding agent capable of specifically binding to the nuclear target to produce more than one nuclear target-associated dsDNA fragments each comprising a first 5' overhang and a second 5' overhang, wherein the transposome comprises a transposase, a first adaptor having the first 5' overhang, and a second adaptor having the second 5' overhang; and

[0058] The more than one nuclear target associated dsDNA fragments or products thereof are barcoded using first one or more oligonucleotide barcodes to produce more than one barcoded nuclear target associated DNA fragments, wherein each of the first one or more oligonucleotide barcodes comprises a first target binding region capable of hybridizing to the more than one nuclear target associated dsDNA fragments or products thereof.

[0059] 18. The method of any one of embodiments 1-17, wherein barcoding the more than one nuclear target-associated dsDNA fragments or products thereof comprises linking the nuclear target-associated dsDNA fragments or products thereof to the first more than one oligonucleotide barcodes.

[0060] 19. The method of any one of embodiments 1-18, wherein barcoding the more than one nuclear target-associated dsDNA fragments or products thereof comprises extending the first more than one oligonucleotide barcodes hybridized to the more than one nuclear target-associated dsDNA fragments or products thereof.

[0061] 20. The method of any one of embodiments 14-19, wherein the first adaptor comprises a first barcode sequence and the second adaptor comprises a second barcode sequence, and optionally the first barcode sequence and / or the second barcode sequence identifies the nuclear target.

[0062] 21. The method according to any one of embodiments 14-20, wherein the first 5' overhang and / or the second 5' overhang comprises a poly(dA) region, a poly(dT) region, or any combination thereof.

[0063] 22. The method according to any one of embodiments 14-21, wherein the first 5' overhang and / or the second 5' overhang comprises a sequence complementary to the first target binding region or a complement thereof.

[0064] 23. The method of any one of embodiments 14-22, wherein the first target binding region comprises a sequence complementary to at least a portion of: a 5' region, a 3' region, or an internal region of a nuclear target-associated dsDNA fragment.

[0065] 24. A method according to any one of embodiments 14-23, wherein the first adaptor and / or the second adaptor comprises a DNA end sequence of the transposon.

[0066] 25. The method of any one of embodiments 1-24, wherein the permeabilization comprises chemical permeabilization or physical permeabilization.

[0067] 26. A method according to any one of embodiments 1-25, wherein the permeabilization comprises contacting the cells with a detergent and / or a surfactant.

[0068] 27. A method according to any one of embodiments 1-26, wherein the permeabilization comprises permeabilizing the cells by sonication.

[0069] 28. A method according to any one of embodiments 1-27, comprising permeabilizing the nuclei in the cells to produce permeabilized nuclei.

[0070] 29. A method according to any one of embodiments 1-28, comprising fixing cells containing the nucleus before permeabilizing the nucleus.

[0071] 30. The method of any one of embodiments 1-29, wherein the nuclear target comprises methylated nucleotides, DNA-associated proteins, chromatin-associated proteins, or any combination thereof.

[0072] 31. The method of any one of embodiments 1-30, wherein the nuclear targets include ALC1, androgen receptor, Bmi-1, BRD4, Brg1, coREST, C-jun, c-Myc, CTCF, EED, EZH2, Fos, histone H1, histone H3, histone H4, heterochromatin protein-1γ, heterochromatin protein-1β, HMGN2 / HMG-17, HP1α, HP1γ, hTERT, Jun, KLF4, K-Ras, Max, MeCP2, MLL / HRX, NPAT, p300, Nanog, NFAT-1, Oct4, P53, Pol II (8WG16), RNA Pol II Ser2P, RNA Pol II Ser5P, RNA Pol II Ser2+5P, RNA Pol II Ser7P, Rb, RNA polymerase II, SMCI, Sox2, STAT1, STAT2, STAT3, Suz12, Tip60, UTF1, or any combination thereof.

[0073] 32. A method according to any one of embodiments 1-31, wherein the binding agent specifically binds to an epitope comprising a methylated (me), phosphorylated (ph), ubiquitinated (ub), para-ubiquitinated (su), biotinylated (bi) or acetylated (ac) histone residue selected from the group consisting of: H1S27ph, H1K25me1, H1K25me2, H1K25me3, H1K26me, H2(A)K4ac, H2(A)K5ac, H2(A)K7 ac, H2(A)S1ph, H2(A)T119ph, H2(A)S122ph, H2(A)S129ph, H2(A)S139ph, H2(A)K119ub, H2(A)K126su, H2(A)K 9bi, H2(A)K13bi, H2(B)K5ac, H2(B)K11ac, H2(B)K12ac, H2(B)K15ac, H2(B)K16ac, H2(B)K20ac, H2(B)S10ph, H2(B)S14ph, H2(B)33ph, H2(B)K120ub, H2(B)K123ub, H3K4ac, H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, H3K56ac, H3K4me1, H3K4me2, H3K4me3, H3R8me, H3K9me1, H3K9me2, H3K9me3, H3R17me, H3K27me1, H3K27me2, H3 K27me3, H3K36me, H3K79me1, H3K79me2, H3K79me3, H3K122ac, H3T3ph, H3S10ph, H3T11ph, H3S28ph, H3K4bi, H3K9bi, H3K18bi, H4K5ac, H4K8ac, H4K12ac, H4K16ac, H4K91ac, H4R3me, H4K20me, H4K59me, H4S1ph, H4K12bi and H4 n-terminal tail ubiquitination or any combination thereof.

[0074] 33. A method according to any one of embodiments 1-32, wherein the binding reagent comprises a tetramer, an aptamer, a protein scaffold or any combination thereof.

[0075] 34. A method according to any one of embodiments 1-33, wherein the binding agent comprises an antibody or a fragment thereof, and optionally the antibody or fragment thereof comprises a monoclonal antibody.

[0076] 35. The method according to embodiment 34, wherein the antibody or fragment thereof comprises Fab, Fab', F(ab') 2, Fv, scFv, dsFv, bispecific antibodies formed by antibody fragments, trispecific antibodies, tetraspecific antibodies, multispecific antibodies, single domain antibodies (sdAb), single chains comprising complementary scFv (tandem scFv) or bispecific tandem scFv, Fv constructs, disulfide-linked Fv, dual variable domain immunoglobulin (DVD-Ig) binding proteins or nanobodies, aptamers, affimers, affimers, affilins, affitins, affimers, alphabodies, anticalins, avimers, DARPins, Fynomers, Kunitz domain peptides, monoclonal antibodies or any combination thereof.

[0077] 36. The method of any one of embodiments 17-35, wherein the binding agent is conjugated to the transposome via chemical coupling, genetic fusion, non-covalent association, or any combination thereof.

[0078] 37. A method according to any one of embodiments 17-36, wherein the conjugate is formed by a 1,3-dipolar cycloaddition reaction, a hetero-Diels-Alder reaction, a nucleophilic substitution reaction, a non-aldol carbonyl reaction, a carbon-carbon multiple bond addition, an oxidation reaction, a click reaction, or any combination thereof.

[0079] 38. A method according to any one of embodiments 17-37, wherein the conjugate is formed by a reaction between acetylene and azide.

[0080] 39. A method according to any one of embodiments 17-38, wherein the conjugate is formed by reaction between an aldehyde or ketone group and a hydrazine or an alkoxyamine.

[0081] 40. The method of any one of embodiments 17-39, wherein the binding agent is conjugated to the transposome via at least one of Protein A, Protein G, Protein A / G, or Protein L.

[0082] 41. The method of any one of embodiments 17-40, wherein the transposase comprises a Tn5 transposase.

[0083] 42. A method according to any one of embodiments 17-41, wherein the transposome comprises a domain that specifically binds to the binding agent, wherein the binding agent and the transposome separate from each other when in contact with the permeabilized cell and enter the cell separately, and wherein the transposome binds to the binding agent within the nucleus of the cell.

[0084] 43. A method according to any one of embodiments 17-41, wherein the transposome is bound to the binding reagent before entering the permeabilized cell, and wherein the transposome bound to the binding reagent is sized to enter the nucleus of the cell through a nuclear pore.

[0085] 44. The method of any one of embodiments 17-43, wherein the conjugate has a diameter of no more than 120 nm.

[0086] 45. The method of any one of embodiments 17-44, wherein the binding reagent and the transposome are each sized to diffuse through nuclear pores of the cell.

[0087] 46. ​​A method according to any one of embodiments 1-45, wherein the permeabilized cells contain an intact nucleus containing chromatin, which remains associated with genomic DNA when the binding reagent is bound to the nuclear target.

[0088] 47. A method according to any one of embodiments 1-46, wherein binding of the binding agent to the nuclear target occurs in the nucleus of the permeabilized cell.

[0089] 48. The method of any one of embodiments 1-47, comprising obtaining sequence data of the more than one barcoded nuclear target-associated DNA fragments or products thereof.

[0090] 49. A method according to embodiment 48, which includes determining information related to the gDNA based on the sequences of the more than one barcoded nuclear target-associated DNA fragments or their products in the sequencing data obtained.

[0091] 50. A method according to embodiment 49, wherein determining the information associated with the gDNA includes determining the genomic information of the gDNA based on the sequences of the more than one barcoded nuclear target-associated DNA fragments in the obtained sequencing data.

[0092] 51. A method according to embodiment 50, comprising digesting nucleosomes associated with the double-stranded gDNA.

[0093] 52. The method of any one of embodiments 49-51, wherein determining the genomic information of the gDNA comprises:

[0094] At least a partial sequence of the gDNA is determined by aligning the sequences of the more than one barcoded nuclear target-associated DNA fragments to a reference sequence of the gDNA.

[0095] 53. A method according to any one of embodiments 49-52, wherein determining the information associated with the gDNA includes determining the methylation group information of the gDNA based on the sequences of the more than one barcoded nuclear target-associated DNA fragments in the obtained sequencing data.

[0096] 54. A method according to embodiment 53, which includes digesting nucleosomes associated with the double-stranded gDNA.

[0097] 55. The method of any one of embodiments 53-54, comprising chemically converting and / or enzymatically converting the cytosine bases of the more than one nuclear target associated dsDNA fragments or products thereof to produce more than one converted nuclear target associated dsDNA fragments having uracil bases, wherein the chemical conversion comprises bisulfite treatment, and wherein the enzymatic conversion comprises APOBEC-mediated conversion.

[0098] 56. The method of embodiment 55, wherein barcoding the more than one nuclear target-associated dsDNA fragments or products thereof comprises barcoding the more than one transformed nuclear target-associated dsDNA fragments or products thereof.

[0099] 57. The method of any one of embodiments 53-56, wherein determining the methylation group information comprises:

[0100] Determine the positions of the more than one barcoded nuclear target-associated DNA fragments in the sequencing data having thymine bases and the corresponding positions of the reference sequence of the gDNA having cytosine bases to determine the corresponding positions of the gDNA having 5-methylcytosine (5mC) bases and / or 5-hydroxymethylcytosine (5hmC) bases.

[0101] 58. A method according to any one of embodiments 1-57, wherein the method comprises Hi-C, chromatin conformation capture (3C), circularized chromatin conformation capture (4C), carbon copy chromosome conformation capture (5C), chromatin immunoprecipitation (ChIP), ChIP-Loop, combined 3C-ChIP-cloning (6C), capture-C or any combination thereof.

[0102] 59. The method of any one of embodiments 1-58, wherein the method comprises Hi-C / ChiP-seq.

[0103] 60. A method according to any one of embodiments 1-59, wherein the method comprises ChiP-seq.

[0104] 61. The method of any one of embodiments 1-60, wherein the cell comprises a copy of a nucleic acid target, the method further comprising:

[0105] contacting a second one or more oligonucleotide barcodes with a copy of the nucleic acid target for hybridization;

[0106] extending the second one or more oligonucleotide barcodes hybridized to the copy of the nucleic acid target to generate one or more barcoded nucleic acid molecules, each of the one or more barcoded nucleic acid molecules comprising a sequence complementary to at least a portion of the nucleic acid target and a molecular tag; and

[0107] Sequence information of more than one barcoded nucleic acid molecule or its product is obtained to determine the copy number of the nucleic acid target in the cell.

[0108] 62. A method according to any one of embodiments 15-61, wherein extending the first one or more oligonucleotide barcodes and / or the second one or more oligonucleotide barcodes comprises extending the one or more oligonucleotide barcodes using a reverse transcriptase and / or a DNA polymerase lacking at least one of a 5' to 3' exonuclease activity and a 3' to 5' exonuclease activity.

[0109] 63. A method according to embodiment 62, wherein the DNA polymerase comprises a Klenow fragment.

[0110] 64. A method according to embodiment 62, wherein the reverse transcriptase comprises a viral reverse transcriptase, optionally wherein the viral reverse transcriptase is murine leukemia virus (MLV) reverse transcriptase or Moloney murine leukemia virus (MMLV) reverse transcriptase.

[0111] 65. A method according to any one of embodiments 61-64, wherein the nucleic acid target comprises a nucleic acid molecule, and optionally the nucleic acid molecule comprises ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNA comprising a poly (A) tail, a sample indexing oligonucleotide, a cellular component binding reagent specific oligonucleotide or any combination thereof.

[0112] 66. A method according to any one of embodiments 1-65, wherein the first target binding region and / or the second target binding region comprises a poly (dA) region, a poly (dT) region, a random sequence, a gene-specific sequence, or any combination thereof.

[0113] 67. A method according to any one of embodiments 48-66, wherein obtaining sequence information of the more than one barcoded nuclear target-associated DNA fragments includes attaching sequencing adapters and / or sequencing primers, their complementary sequences and / or portions thereof to the more than one barcoded nuclear target-associated DNA fragments or their products.

[0114] 68. A method according to any one of embodiments 61-67, wherein obtaining sequence information of the more than one barcoded nucleic acid molecules includes attaching sequencing adapters and / or sequencing primers, their complementary sequences and / or portions thereof to the more than one barcoded nucleic acid molecules or their products.

[0115] 69. The method of any one of embodiments 1-68, wherein each of the first more than one oligonucleotide barcodes comprises a first universal sequence, and wherein each of the second more than one oligonucleotide barcodes comprises a second universal sequence.

[0116] 70. A method according to embodiment 69, wherein the first universal sequence and the second universal sequence are the same.

[0117] 71. A method according to embodiment 69, wherein the first universal sequence and the second universal sequence are different.

[0118] 72. A method according to any one of embodiments 69-71, wherein the first universal sequence and / or the second universal sequence comprises a binding site for: a sequencing primer and / or a sequencing adapter, a complementary sequence thereof and / or a portion thereof.

[0119] 73. A method according to any one of embodiments 16-72, wherein the sequencing adapter comprises a P5 sequence, a P7 sequence, a complementary sequence thereof and / or a portion thereof.

[0120] 74. A method according to any one of embodiments 16-73, wherein the sequencing primers include a read segment 1 sequencing primer, a read segment 2 sequencing primer, their complementary sequences and / or portions thereof.

[0121] 75. The method of any one of embodiments 1-74, wherein the first one or more oligonucleotide barcodes and the second one or more oligonucleotide barcodes each comprise a molecular tag.

[0122] 76. The method of any one of embodiments 1-74, wherein at least 10 of the first more than one oligonucleotide barcodes and the second more than one oligonucleotide barcodes comprise different molecular marker sequences.

[0123] 77. A method according to any one of embodiments 1-76, wherein each molecular label of the first more than one oligonucleotide barcode and the second more than one oligonucleotide barcode comprises at least 6 nucleotides.

[0124] 78. The method of any one of embodiments 1-77, wherein the first one or more oligonucleotide barcodes and the second one or more oligonucleotide barcodes are associated with a solid support.

[0125] 79. The method of embodiment 78, wherein the first one or more oligonucleotide barcodes and the second one or more oligonucleotide barcodes associated with the same solid support each comprise the same sample label.

[0126] 80. A method according to embodiment 79, wherein each sample label of the first more than one oligonucleotide barcode and the second more than one oligonucleotide barcode comprises at least 6 nucleotides.

[0127] 81. A method according to any one of embodiments 1-80, wherein the first more than one oligonucleotide barcode and the second more than one oligonucleotide barcode each comprise a cell marker, and optionally each cell marker of the first more than one oligonucleotide barcode and the second more than one oligonucleotide barcode comprises at least 6 nucleotides.

[0128] 82. A method according to embodiment 81, wherein the oligonucleotide barcodes of the first more than one oligonucleotide barcodes and the second more than one oligonucleotide barcodes associated with the same solid support comprise the same cellular marker.

[0129] 83. The method of any one of embodiments 81-82, wherein the oligonucleotide barcodes in the first one or more oligonucleotide barcodes and the second one or more oligonucleotide barcodes associated with different solid supports comprise different cellular markers.

[0130] 84. The method of any one of embodiments 78-83, wherein the solid support comprises synthetic particles, a flat surface, or a combination thereof.

[0131] 85. The method of any one of embodiments 1-84, comprising associating a synthetic particle comprising the first more than one oligonucleotide barcode and the second more than one oligonucleotide barcode with the cell.

[0132] 86. The method of embodiment 85, comprising lysing the cells after associating the synthetic particles with the cells, optionally lysing the cells comprising heating the cells, contacting the cells with a detergent, changing the pH of the cells, or any combination thereof.

[0133] 87. A method according to any one of embodiments 85-86, wherein the synthetic particles and the single cells are in the same partition, and optionally the partition is a well or a droplet.

[0134] 88. A method according to any one of embodiments 78-87, wherein at least one of the first more than one oligonucleotide barcodes and the second more than one oligonucleotide barcodes is fixed or partially fixed on the synthetic particle, or at least one of the first more than one oligonucleotide barcodes and the second more than one oligonucleotide barcodes is encapsulated or partially encapsulated in the synthetic particle.

[0135] 89. The method of any one of embodiments 78-88, wherein the synthetic particles are destructible, optionally the synthetic particles are destructible hydrogel particles.

[0136] 90. A method according to any one of embodiments 78-89, wherein the synthetic particles comprise beads, optionally the beads comprise agarose gel beads, streptavidin beads, agarose beads, magnetic beads, conjugated beads, protein A conjugated beads, protein G conjugated beads, protein A / G conjugated beads, protein L conjugated beads, oligo (dT) conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads or any combination thereof.

[0137] 91. A method according to any one of embodiments 78-90, wherein the synthetic particles comprise a material selected from the group consisting of polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic substances, ceramics, plastics, glass, methylstyrene, acrylic polymers, titanium, latex, agarose gel, cellulose, nylon, silicone, and any combination thereof.

[0138] 92. The method according to any one of embodiments 78-91,

[0139] wherein each of the first one or more oligonucleotide barcodes and the second one or more oligonucleotide barcodes comprises a linker functional group,

[0140] wherein the synthetic particles comprise solid support functional groups, and

[0141] wherein the support functional group and the linker functional group are associated with each other, and optionally the linker functional group and the support functional group are individually selected from the group consisting of C6, biotin, streptavidin, one or more primary amines, one or more aldehydes, one or more ketones, and any combination thereof.

[0142] 93. A kit comprising:

[0143] A digestion composition comprising a DNA digestion enzyme and a binding agent capable of specifically binding to a nuclear target, wherein the binding agent comprises a binding agent-specific oligonucleotide comprising a unique identifier sequence for the binding agent.

[0144] 94. A kit according to embodiment 93, wherein the DNA digestion enzyme comprises a restriction enzyme, micrococcal nuclease I, a transposase, a functional fragment thereof or any combination thereof, optionally the transposase comprises Tn5 transposase, and further optionally the digestion composition comprises at least one of a first adaptor having a first 5' overhang and a second adaptor having a second 5' overhang.

[0145] 95. A kit according to any one of embodiments 93-94, wherein the digestion composition comprises a fusion protein comprising the binding agent and the digestive enzyme, optionally wherein the fusion protein is sized to diffuse through the nuclear pores of a cell.

[0146] 96. A kit according to any one of embodiments 93-94, wherein the digestive composition comprises a conjugate comprising the binding agent and the digestive enzyme, optionally wherein the conjugate is sized to diffuse through the nuclear pores of a cell.

[0147] 97. A kit according to any one of embodiments 93-94, wherein the DNA digesting enzyme comprises a domain capable of specifically binding to the binding reagent, optionally the domain of the DNA digesting enzyme comprises at least one of protein A, protein G, protein A / G or protein L, and further optionally the size of the DNA digesting enzyme bound to the binding reagent is configured to diffuse through the nuclear pores of cells.

[0148] 98. The kit of any one of embodiments 93-97, further comprising a protease, optionally comprising proteinase K.

[0149] 99. A kit comprising:

[0150] A conjugate comprising a transposome and a binding agent capable of specifically binding to a nuclear target, wherein the transposome comprises a transposase, a first adaptor having a first 5' overhang, and a second adaptor having a second 5' overhang.

[0151] 100. The kit of any one of embodiments 93-99, further comprising a DNA polymerase lacking at least one of a 5' to 3' exonuclease activity and a 3' to 5' exonuclease activity, optionally wherein the DNA polymerase comprises a Klenow fragment.

[0152] 101. The kit of any one of embodiments 93-100, further comprising a reverse transcriptase, optionally wherein the reverse transcriptase comprises a viral reverse transcriptase, and optionally wherein the viral reverse transcriptase is a murine leukemia virus (MLV) reverse transcriptase or a Moloney murine leukemia virus (MMLV) reverse transcriptase.

[0153] 102. The kit of any one of embodiments 93-101, further comprising one or more of a ligase, a detergent, a surfactant, a buffer, and a cartridge.

[0154] 103. A kit according to any one of embodiments 93-102, wherein the kit comprises one or more reagents for reverse transcription reaction and / or amplification reaction.

[0155] 104. A kit according to any one of embodiments 93-103, wherein the nuclear target comprises a DNA-associated protein or a chromatin-associated protein.

[0156] 105. A kit according to any of embodiments 93-104, wherein the nuclear target comprises methylated nucleotides.

[0157] 106. A kit according to any one of embodiments 93-105, wherein the nuclear targets include ALC1, androgen receptor, Bmi-1, BRD4, Brg1, coREST, C-jun, c-Myc, CTCF, EED, EZH2, Fos, histone H1, histone H3, histone H4, heterochromatin protein-1γ, heterochromatin protein-1β, HMGN2 / HMG-17, HP1α, HP1γ, hTERT, Jun, KLF4, K-Ras, Max, MeCP2, MLL / HRX, NPAT, p300, Nanog, NFAT-1, Oct4, P53, Pol II (8WG16), RNAPol II Ser2P, RNA Pol II Ser5P, RNA Pol II Ser2+5P, RNAPol II Ser7P, Rb, RNA polymerase II, SMCI, Sox2, STAT1, STAT2, STAT3, Suz12, Tip60, UTF1, or any combination thereof.

[0158] 107. A kit according to any one of embodiments 93-106, wherein the binding agent specifically binds to an epitope comprising a methylated (me), phosphorylated (ph), ubiquitinated (ub), para-ubiquitinated (su), biotinylated (bi) or acetylated (ac) histone residue selected from the group consisting of: H1S27ph, H1K25me1, H1K25me2, H1K25me3, H1K26me, H2(A)K4ac, H2(A)K5ac, H2(A)K6ac, H2(A)K7ac, H2(A)K8ac, H2(A)K9ac, H2(A)K10ac, H2(A)K11ac, H2(A)K12ac, H2(A)K13ac, H2(A)K14ac, H2(A)K15ac, H2(A)K16ac, H2(A)K17ac, H2(A)K18ac, H2(A)K19ac, H2(A)K22ac, H2(A)K23ac, H2(A)K24ac, H2(A)K25 )K7ac, H2(A)S1ph, H2(A)T119ph, H2(A)S122ph, H2(A)S129ph, H2(A)S139ph, H2(A)K119ub, H2(A)K126su, H2( A)K9bi, H2(A)K13bi, H2(B)K5ac, H2(B)K11ac, H2(B)K12ac, H2(B)K15ac, H2(B)K16ac, H2(B)K20ac, H2(B)S10p h, H2(B)S14ph, H2(B)33ph, H2(B)K120ub, H2(B)K123ub, H3K4ac, H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27a c. H3K56ac, H3K4me1, H3K4me2, H3K4me3, H3R8me, H3K9me1, H3K9me2, H3K9me3, H3R17me, H3K27me1, H3K27me2, H 3K27me3, H3K36me, H3K79me1, H3K79me2, H3K79me3, H3K122ac, H3T3ph, H3S10ph, H3T11ph, H3S28ph, H3K4bi, H3K9bi, H3K18bi, H4K5ac, H4K8ac, H4K12ac, H4K16ac, H4K91ac, H4R3me, H4K20me, H4K59me, H4S1ph, H4K12bi and H4 n-terminal tail ubiquitination or any combination thereof.

[0159] 108. A kit according to any one of embodiments 93-107, wherein the binding reagent comprises a tetramer, an aptamer, a protein scaffold or any combination thereof.

[0160] 109. A kit according to any one of embodiments 93-108, wherein the binding reagent comprises an antibody or fragment thereof, optionally the antibody or fragment thereof comprises a monoclonal antibody.

[0161] 110. The kit according to embodiment 109, wherein the antibody or fragment thereof comprises Fab, Fab', F(ab')2 , Fv, scFv, dsFv, bispecific antibodies formed by antibody fragments, trispecific antibodies, tetraspecific antibodies, multispecific antibodies, single domain antibodies (sdAb), single chains comprising complementary scFv (tandem scFv) or bispecific tandem scFv, Fv constructs, disulfide-linked Fv, dual variable domain immunoglobulin (DVD-Ig) binding proteins or nanobodies, aptamers, affimers, affimers, affilins, affitins, affimers, alphabodies, anticalins, avimers, DARPins, Fynomers, Kunitz domain peptides, monoclonal antibodies or any combination thereof.

[0162] 111. A kit according to any one of embodiments 99-110, wherein the binding reagent is conjugated to the transposome via chemical coupling, genetic fusion, non-covalent association, or any combination thereof.

[0163] 112. A kit according to any one of embodiments 99-111, wherein the conjugate is formed by a 1,3-dipolar cycloaddition reaction, a hetero-Diels-Alder reaction, a nucleophilic substitution reaction, a non-aldol carbonyl reaction, a carbon-carbon multiple bond addition, an oxidation reaction, a click reaction, or any combination thereof.

[0164] 113. A kit according to any one of embodiments 99-112, wherein the conjugate is formed by a reaction between acetylene and azide.

[0165] 114. A kit according to any one of embodiments 99-113, wherein the conjugate is formed by reaction between an aldehyde or ketone group and a hydrazine or an alkoxyamine.

[0166] 115. A kit according to any one of embodiments 99-114, wherein the binding reagent is conjugated to the transposome via at least one of Protein A, Protein G, Protein A / G, or Protein L.

[0167] 116. A kit according to any one of embodiments 99-115, wherein the transposase comprises Tn5 transposase.

[0168] 117. A kit according to any one of embodiments 99-116, wherein the conjugate has a diameter of no more than 120 nm.

[0169] 118. A kit according to any one of embodiments 99-117, wherein the binding reagent and the transposome are each sized to diffuse through nuclear pores.

[0170] 119. A kit according to any one of embodiments 93-118, wherein the kit comprises more than one oligonucleotide barcode, wherein each of the more than one oligonucleotide barcodes comprises a molecular marker and a target binding region, and wherein at least 10 of the more than one oligonucleotide barcodes comprise different molecular marker sequences.

[0171] 120. The kit of embodiment 119, wherein the target binding region comprises a gene-specific sequence, an oligo(dT) sequence, a random polymer, or any combination thereof.

[0172] 121. A kit according to any one of embodiments 119-120, wherein the oligonucleotide barcodes comprise the same sample marker and / or the same cell marker, and optionally each sample marker and / or cell marker of the more than one oligonucleotide barcode comprises at least 6 nucleotides.

[0173] 122. A kit according to any one of embodiments 119-121, wherein each molecular label of the more than one oligonucleotide barcode comprises at least 6 nucleotides.

[0174] 123. A kit according to any one of embodiments 119-122, wherein at least one of the more than one oligonucleotide barcodes is partially immobilized on the synthetic particle, wherein at least one of the more than one oligonucleotide barcodes is immobilized or partially immobilized on the synthetic particle; and / or at least one of the more than one oligonucleotide barcodes is encapsulated or partially encapsulated in the synthetic particle.

[0175] 124. A kit according to embodiment 123, wherein the synthetic particles comprise beads, optionally wherein the beads comprise agarose gel beads, streptavidin beads, agarose beads, magnetic beads, conjugated beads, protein A conjugated beads, protein G conjugated beads, protein A / G conjugated beads, protein L conjugated beads, oligo (dT) conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads or any combination thereof.

[0176] 125. A kit according to any of embodiments 123-124, wherein the synthetic particles comprise a material selected from the group consisting of polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic substances, ceramics, plastics, glass, methylstyrene, acrylic polymers, titanium, latex, agarose gel, cellulose, nylon, silicone, and any combination thereof.

[0177] 126. The kit according to any one of embodiments 123-125,

[0178] wherein each of the more than one oligonucleotide barcodes comprises a linker functional group,

[0179] wherein the synthetic particles comprise solid support functional groups, and

[0180] wherein the support functional group and the linker functional group are associated with each other.

[0181] 127. A kit according to embodiment 126, wherein the linker functional group and the support functional group are individually selected from the group consisting of C6, biotin, streptavidin, one or more primary amines, one or more aldehydes, one or more ketones and any combination thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0183] Figure 1 is a schematic diagram illustrating a binding agent (eg, a protein binding agent) that binds a fusion protein comprising a DNA digesting enzyme according to some embodiments herein.

[0184] Figure 2 is a diagram illustrating the permeabilization of cells with Figure 1 Schematic diagram of contacting a binding agent (e.g., a protein binding agent) that binds to a fusion protein such that the protein binding agent that binds to the fusion protein binds to a target protein associated with DNA in a cell nucleus.

[0185] Figure 3 is a diagram showing the digestion of DNA with a DNA digestion enzyme to produce Figure 1 Schematic diagram of a complex of a binding agent (eg, protein binding agent) that binds to a fusion protein and digested DNA. According to some embodiments herein, the complex can be separated (eg, removed) from the cell.

[0186] Figure 4 is a schematic diagram illustrating a substrate (eg, a bead) comprising a first oligonucleotide probe and a second oligonucleotide probe immobilized on the substrate according to some embodiments herein.

[0187] Figures 5A-5D is a schematic diagram illustrating the capture and processing of DNA from single cells. Figure 5A is a diagram illustrating the use of some embodiments of the present invention Figure 4 The substrate (e.g., beads) captures Figure 3 Schematic diagram of the complex. Figure 5B is a diagram illustrating some embodiments according to the present invention Figure 5A Schematic diagram of oligonucleotide ligation and reverse transcription of the capture complex. Figure 5C Illustrated is the digestion of the protein binding reagent, fusion protein, and target protein to produce beads associated only with oligonucleotides according to some embodiments herein. Figure 5DDenaturation, random priming and extension, and polymerase chain reaction to generate a library according to some embodiments herein are illustrated.

[0188] Figure 6 is a schematic diagram illustrating the generation of a library of target protein-associated DNA according to some embodiments herein. Generating the library can include providing components of oligonucleotides that bind to a substrate.

[0189] Figures 7A-7C is a schematic representation of labeling an oligonucleotide probe with a target binding region nucleic acid sequence according to some embodiments herein. Fig. 7A Illustration of beads with oligonucleotide probes having poly-T hybridization domains according to some embodiments herein. Figure 7B A double-stranded DNA (dsDNA) having a template strand sequence (comprising a poly-A hybridization domain and a target nucleic acid sequence) and a complementary strand sequence (comprising a target binding region complementary to the target nucleic acid sequence) is illustrated, wherein the dsDNA is attached to biotin according to some embodiments herein. Figure 7C Some embodiments according to the present invention are shown in FIG. Fig. 7A The oligonucleotide probes Figure 7B Hybridization of dsDNA.

[0190] Figure 8 Schematically illustrating that according to some embodiments herein, after the complementary strand sequence of the dsDNA is ligated to the oligonucleotide probe, the dsDNA is denatured and the template strand sequence is removed by binding of biotin to streptavidin.

[0191] Fig. 9 Schematically illustrates some embodiments according to the present invention, such as Figure 8 The oligonucleotide probe having a complementary target strand sequence is shown binding to the target nucleic acid sequence.

[0192] Fig. 10A -B is a non-limiting schematic representation of a conjugate provided herein for labeling nuclear target-associated DNA from a single cell.

[0193] Fig.11 Non-limiting exemplary bar codes are illustrated.

[0194] Fig.12 A non-limiting exemplary workflow for barcoding and digital counting is shown.

[0195] Fig.13 is a schematic representation showing a non-limiting exemplary process for generating an indexed library of targets barcoded at the 3' end from more than one target.

[0196] Details

[0197] Reference is made to the accompanying drawings forming a part of this document in the following detailed description. In the accompanying drawings, similar symbols generally identify similar components unless the context otherwise indicates. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized and other changes may be made without departing from the spirit or scope of the subject matter provided herein. It is readily understood that the aspects of the present disclosure as generally described herein and illustrated in the drawings may be arranged, replaced, combined, separated, and designed in a variety of different configurations, all of which are expressly contemplated herein and constitute a part of the disclosure herein.

[0198] All patents, published patent applications, other publications, and sequences from GenBank and other databases cited herein that relate to the related art are incorporated by reference in their entirety.

[0199] Quantification of small numbers of nucleic acids (e.g., messenger ribonucleic acid (mRNA) molecules) is clinically important for determining genes expressed in cells, for example, at different developmental stages or under different environmental conditions. However, determining the absolute number of nucleic acid molecules (e.g., mRNA molecules) can also be very challenging, especially when the number of molecules is very small. One method for determining the absolute number of molecules in a sample is the digital polymerase chain reaction (PCR). Ideally, PCR produces identical copies of molecules in each cycle. However, PCR may have the disadvantage that each molecule is replicated with a random probability, and this probability varies depending on the PCR cycle and the gene sequence, which leads to amplification bias and inaccurate gene expression measurements. Random barcodes with unique molecular labels (molecular labels, also called molecular indexes (MI)) can be used to count the number of molecules and correct for amplification bias. Random barcoding such as Precise TM Assay (Cellular Research, Inc. (Palo Alto, CA)) and Rhapsody TM The PCR amplification assay (Becton, Dickinson and Company (Franklin Lakes, NJ)) can correct for bias induced by PCR and library preparation steps by labeling mRNA during reverse transcription (RT) using a molecular marker (ML).

[0200] Precise TMThe assay can utilize a non-exhaustive pool of random barcodes with unique molecular marker sequences on a large number (e.g., 6561 to 65536) of poly (T) oligonucleotides, which are hybridized with all poly (A) mRNAs in the sample during the RT step. The random barcodes can include universal PCR priming sites. During RT, the target gene molecules react randomly with the random barcodes. Each target molecule can hybridize with the random barcodes, resulting in the generation of randomly barcoded complementary ribonucleotides (cDNA) molecules. After labeling, the randomly barcoded cDNA molecules from the microwells of the microplate can be pooled into a single tube for PCR amplification and sequencing. The raw sequencing data can be analyzed to generate the number of reads, the number of random barcodes with unique molecular marker sequences, and the number of mRNA molecules.

[0201] According to some embodiments herein, methods, kits and compositions for marking DNA associated with target proteins from single cells are described. In the methods, kits and compositions of some embodiments, the first composition comprises a binding agent (e.g., a protein binding agent), is substantially composed of a binding agent (e.g., a protein binding agent) or is composed of a binding agent (e.g., a protein binding agent), which specifically binds to a target protein (e.g., chromatin) associated with DNA in a cell, thereby forming a complex. The binding agent (e.g., a protein binding agent) may also comprise a unique identifier such as a nucleic acid barcode. The second composition may comprise a fusion protein combined with a protein binding agent, is substantially composed of a fusion protein combined with a protein binding agent or is composed of a fusion protein combined with a protein binding agent. The fusion protein may comprise a DNA digestion enzyme that digests the DNA associated with the target protein, producing a complex comprising a protein binding agent containing a unique identifier, a fusion protein and the digested DNA. Then barcoding may be performed to mark the specific digested DNA associated with a specific unique identifier of the protein binding agent. Therefore, a specific DNA sequence associated with a specific target protein may be identified at the single cell and single molecule level. For example, a barcode-containing bead can be contacted with a protein binding reagent, each of which shares a common sample identifier sequence, and the protein binding reagent is compounded with the digested DNA so that the digested DNA and the unique identifier can each be associated with the same sample identifier sequence. In addition, in order to determine an accurate description of the molecular interaction in the cell, the complex can be formed in situ in the permeabilized cell or nucleus (and the DNA can be digested). Not limited by theory, it is considered that the in situ formation of the complex can avoid artifacts (artifactual) and false positive interactions that may be produced by labeling cell extracts or their fractions. Optionally, multiple analysis can be performed to identify the DNA associated with two or more different types of target proteins in the same sample (for example, in the same single cell). Some embodiments described herein relate to methods for DNA labeling of target protein association in single cells using kits and / or compositions described herein. Some embodiments described herein relate to kits and / or compositions.

[0202] Conventional methods for labeling DNA associated with target proteins can include methods such as cross-linking the target protein to the associated DNA, DNA fragmentation, immunoprecipitation, DNA separation and sequencing. Such methods generally involve the analysis of thousands or millions of cells to label and analyze the DNA associated with the target protein. Therefore, such methods are not suitable for analyzing the DNA associated with the target protein in a single cell. The kits, methods and compositions of some embodiments described herein produce accurate and repeatable identification and analysis of the DNA associated with the target protein in a single cell. In some embodiments, the target protein associated with the DNA in a single cell can be identified in situ (e.g., in the nucleus of the cell), thereby identifying the target protein-DNA interaction, which accurately describes the interaction occurring in the cell itself.

[0203] The disclosure herein includes methods for labeling nuclear target associated DNA in a cell. In some embodiments, the method includes: permeabilizing a cell comprising a nuclear target associated with a double-stranded deoxyribonucleic acid (dsDNA). The dsDNA may be genomic DNA (gDNA). The method may include: contacting the nuclear target with a conjugate to produce more than one nuclear target associated dsDNA fragments (e.g., nuclear target associated gDNA fragments) each comprising a first 5' overhang and a second 5' overhang, the conjugate comprising a transposome and a binding agent capable of specifically binding to a nuclear target, wherein the transposome comprises a transposase, a first adaptor having a first 5' overhang, and a second adaptor having a second 5' overhang; and barcoding the more than one nuclear target associated dsDNA fragments or products thereof using a first more than one oligonucleotide barcode to produce more than one barcoded nuclear target associated DNA fragments, wherein each of the first more than one oligonucleotide barcodes comprises a first target binding region capable of hybridizing with the more than one nuclear target associated dsDNA fragments or products thereof. Barcoded nuclear target-associated DNA fragments can be single-stranded or double-stranded.

[0204] The disclosure herein includes a kit. In some embodiments, the kit comprises: a conjugate comprising a transposome and a binding agent capable of specifically binding to a nuclear target, wherein the transposome comprises a transposase, a first adapter having a first 5' overhang, and a second adapter having a second 5' overhang. The kit may comprise: a DNA polymerase lacking at least one of a 5' to 3' exonuclease activity and a 3' to 5' exonuclease activity, optionally wherein the DNA polymerase comprises a Klenow fragment. The kit may comprise: a reverse transcriptase, optionally wherein the reverse transcriptase comprises a viral reverse transcriptase, and optionally the viral reverse transcriptase is a murine leukemia virus (MLV) reverse transcriptase or a Moloney murine leukemia virus (MMLV) reverse transcriptase. The kit may comprise: a ligase. The kit may comprise: a detergent and / or a surfactant. The kit may comprise: a buffer, a cartridge, or both. The kit may comprise: one or more reagents for reverse transcription reactions and / or amplification reactions.

[0205] The disclosure herein includes methods of labeling nuclear target associated DNA in a cell. In some embodiments, the method includes: permeabilizing a cell comprising a nuclear target associated with double-stranded deoxyribonucleic acid (dsDNA), wherein optionally the dsDNA is genomic DNA (gDNA). The method may include: contacting the nuclear target with a digestion composition to produce more than one nuclear target associated dsDNA fragments (e.g., nuclear target associated gDNA fragments) each comprising a single-stranded overhang, the digestion composition comprising a DNA digestion enzyme and a binding agent capable of specifically binding to a nuclear target, wherein each binding agent comprises a binding agent-specific oligonucleotide comprising a unique identifier sequence for the binding agent. The method may include: barcoding more than one nuclear target associated dsDNA fragments or products thereof using first one or more oligonucleotide barcodes to produce more than one barcoded nuclear target associated DNA fragments, each of the more than one barcoded nuclear target associated DNA fragments comprising a sequence complementary to at least a portion of the nuclear target associated dsDNA fragments, wherein each of the first one or more oligonucleotide barcodes comprises a first target binding region capable of hybridizing to the more than one nuclear target associated dsDNA fragments or products thereof. The barcoded nuclear target associated DNA fragments may be single-stranded or double-stranded. The method may include: barcoding a binding reagent-specific oligonucleotide or a product thereof using a second one or more oligonucleotide barcodes to produce more than one barcoded binding reagent-specific oligonucleotides, each of the more than one barcoded binding reagent-specific oligonucleotides comprising a sequence complementary to at least a portion of a unique identifier sequence, wherein each of the second one or more oligonucleotide barcodes comprises a second target binding region capable of hybridizing to the binding reagent-specific oligonucleotide or a product thereof.

[0206] The disclosure herein includes kits. The kits can include a digestion composition comprising a DNA digestion enzyme and a binding agent capable of specifically binding to a nuclear target, wherein the binding agent comprises a binding agent-specific oligonucleotide comprising a unique identifier sequence for the binding agent.

[0207] definition

[0208] Unless otherwise defined, the technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure belongs. See, for example, Singleton et al., Dictionary of Microbiology and Molecular Biology, 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989). For the purposes of the present disclosure, the following terms are defined below.

[0209] As used herein, the term "adapter" may mean a sequence that promotes the amplification or sequencing of an associated nucleic acid. The associated nucleic acid may include a target nucleic acid. The associated nucleic acid may include one or more spatial markers, target markers, sample markers, index markers or barcode sequences (e.g., molecular markers). The adaptor may be linear. The adaptor may be a pre-adenylated adaptor. The adaptor may be double-stranded or single-stranded. One or more adaptors may be located at the 5' end or 3' end of the nucleic acid. When the adaptor comprises a known sequence at the 5' end and the 3' end, the known sequence may be the same or different sequences. The adaptor located at the 5' end and / or the 3' end of the polynucleotide may be able to hybridize with one or more oligonucleotides fixed on the surface. In some embodiments, the adaptor may include a universal sequence. The universal sequence may be a region of nucleotide sequences common to two or more nucleic acid molecules. Two or more nucleic acid molecules may also have regions of different sequences. Therefore, for example, the 5' adaptor may include the same and / or universal nucleic acid sequence, and the 3' adaptor may include the same and / or universal sequence. The universal sequence that can be present in different members of more than one nucleic acid molecule can allow the use of a single universal primer complementary to the universal sequence to replicate or amplify more than one different sequence. Similarly, at least one, two (e.g., a pair) or more universal sequences that can be present in different members of a set of nucleic acid molecules can allow the use of at least one, two (e.g., a pair) or more single universal primers complementary to the universal sequence to replicate or amplify more than one different sequence. Therefore, the universal primer comprises a sequence that can hybridize with such universal sequences. The molecule carrying the target nucleic acid sequence can be modified to attach a universal adapter (e.g., a non-target nucleic acid sequence) to one or both ends of different target nucleic acid sequences. One or more universal primers attached to the target nucleic acid can provide a site for universal primer hybridization. One or more universal primers attached to the target nucleic acid can be identical or different from each other.

[0210] As used herein, the term "association" or "associated with..." can mean that two or more substances can be identified as being co-located at a certain point in time. Association can mean that two or more substances are or have been in a similar container. Association can be an informatics association. For example, digital information about two or more substances can be stored and can be used to determine that one or more substances are co-located at a certain point in time. Association can also be a physical association. In some embodiments, two or more associated substances are "tethered", "attached" or "fixed" to each other or to a common solid or semi-solid surface. Association can refer to a covalent or non-covalent means for attaching a label to a solid or semi-solid support (such as a bead). Association can be a covalent bond between a target and a label. Association can include hybridization between two molecules (such as a target molecule and a label).

[0211] As used herein, the term "complementary" can refer to the ability of accurate pairing between two nucleotides. For example, if the nucleotide at a given position of a nucleic acid can hydrogen bond with the nucleotide of another nucleic acid, the two nucleic acids are considered to be complementary to each other at that position. The complementarity between two single-stranded nucleic acid molecules can be "partial", where only some nucleotides in the nucleotides bind, or when there is complete complementarity between single-stranded molecules, this complementarity can be complete. If the first nucleotide sequence is complementary to the second nucleotide sequence, the first nucleotide sequence can be referred to as the "complement" of the second sequence. If the first nucleotide sequence is complementary to the reverse sequence of the second sequence (that is, the nucleotide order is opposite), the first nucleotide sequence can be referred to as the "reverse complement" of the second sequence. As used herein, a "complementary" sequence can refer to the "complement" or "reverse complement" of a sequence. It is understood from the present disclosure that if a molecule can hybridize with another molecule, it can be complementary or partially complementary to the molecule it hybridizes with.

[0212] As used herein, the term "digital counting" may refer to a method for estimating the number of target molecules in a sample. Digital counting may include the step of determining the number of unique markers that have been associated with the target in the sample. This method, which may be stochastic in nature, converts the problem of counting molecules from one of localization and identification of identical molecules to a series of yes / no digital questions about detecting a set of predefined markers.

[0213] As used herein, the term "a label" or "more than one label" may refer to a nucleic acid code associated with a target in a sample. A label may be, for example, a nucleic acid label. A label may be a fully or partially amplifiable label. A label may be a fully or partially sequenceable label. A label may be a portion of a natural nucleic acid that is identifiable as being distinct. A label may be a known sequence. A label may include a junction of a nucleic acid sequence, such as a junction of a natural and a non-natural sequence. As used herein, the term "label" may be used interchangeably with the term "index," "tag," or "label-tag." A label may convey information. For example, in various embodiments, a label may be used to determine the identity of a sample, the source of a sample, the identity of a cell, and / or a target.

[0214] As used herein, the term "non-depleting reservoir" may refer to a pool of barcodes (e.g., random barcodes) composed of many different markers. A non-depleting reservoir may include a large number of different barcodes such that when the non-depleting reservoir is associated with a target pool, each target may be associated with a unique barcode. The uniqueness of each labeled target molecule can be determined by the statistics of random selection and depends on the number of copies of the same target molecule in the set compared to the diversity of the marker. The size of the resulting set of labeled target molecules can be determined by the stochastic nature of the barcoding process, and then analysis of the number of detected barcodes allows calculation of the number of target molecules present in the original set or sample. When the ratio of the number of copies of the target molecules present to the number of unique barcodes is low, the labeled target molecules are highly unique (i.e., the probability that more than one target molecule is labeled by a given marker is very low).

[0215] Various nucleic acids are described according to some embodiments herein. For example, an oligonucleotide, a sample and / or a target can include a nucleic acid.

[0216] As used herein, the term "nucleic acid" refers to a polynucleotide sequence or a fragment thereof. Nucleic acids may include nucleotides. Nucleic acids may be exogenous or endogenous to a cell. Nucleic acids may be present in a cell-free environment. Nucleic acids may be genes or fragments thereof. Nucleic acids may be DNA. Nucleic acids may be RNA. Nucleic acids may include one or more analogs (e.g., altered backbones, sugars, or nucleic acid bases). Some non-limiting examples of analogs include: 5-bromouracil, peptide nucleic acids, heterologous nucleic acids, morpholinos, locked nucleic acids, ethylene glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein linked to a sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queuosine, and wyosine. "Nucleic acid," "polynucleotide," "target polynucleotide," and "target nucleic acid" may be used interchangeably.

[0217] As used herein, "nucleoside" has the common and customary meanings understood by those of ordinary skill in the art according to this specification, and may include natural nucleosides, such as 2'-deoxy and 2'-hydroxyl forms. "Analogs" about nucleosides may include synthetic nucleosides containing modified base moieties and / or modified sugar moieties, etc. Analogs may be able to hybridize. Analogs may include synthetic nucleosides designed to enhance binding properties, reduce complexity, increase specificity, etc. Exemplary types of analogs may include oligonucleotide phosphoramidates (referred to herein as "amide esters"), peptide nucleic acids (referred to herein as "PNA"), oligo-2'-O-alkyl ribonucleotides, polynucleotides containing C-5 propynyl pyrimidines, and locked nucleic acids (LNA).

[0218] As used herein, "upstream" (and variants of this root term) has the ordinary and customary meaning as understood by those of ordinary skill in the art in light of this specification, and refers to a position on a nucleic acid relative to 5' (e.g., 5' compared to a reference position). As used herein, "downstream" (and variants of this root term) has the ordinary and customary meaning as understood by those of ordinary skill in the art in light of this specification, and refers to a position on a nucleic acid relative to 3' (e.g., 3' compared to a reference position).

[0219] Nucleic acids may include one or more modifications (e.g., base modifications, backbone modifications) to provide new or enhanced features (e.g., improved stability) to nucleic acids. Nucleic acids may include nucleic acid affinity tags. Nucleosides may be base-sugar combinations. The base portion of a nucleoside may be a heterocyclic base. The two most common categories of such heterocyclic bases are purine and pyrimidine. Nucleotides may be nucleosides further comprising a phosphate group covalently linked to the sugar portion of the nucleoside. For those nucleosides including furanose, the phosphate group may be linked to the 2', 3' or 5' hydroxyl portion of the sugar. In forming nucleic acids, the phosphate group may covalently link adjacent nucleosides to each other to form a linear polymer compound. Subsequently, each end of this linear polymer compound may be further linked to form a cyclic compound; however, linear compounds are generally suitable. In addition, linear compounds may have internal nucleotide base complementarity, and may therefore be folded in a manner that produces a fully or partially double-stranded compound. In nucleic acids, a phosphate group may generally refer to a skeleton between nucleosides forming a nucleic acid. A bond or skeleton may be a 3' to 5' phosphodiester bond.

[0220] Nucleic acid can include modified backbone and / or modified internucleoside bond.Modified backbone can include those backbones that retain phosphorus atom in the backbone and those backbones that do not have phosphorus atom in the backbone.Suitable wherein phosphorus atom-containing modified nucleic acid backbone can include, for example, phosphorothioate, chiral phosphorothioate, phosphorodithioate, phosphotriester, aminoalkylphosphotriester, methyl and other alkylphosphonates such as 3'-alkylenephosphonates, 5'-alkylenephosphonates, chiral phosphonates, phosphites, phosphoramidates include 3'-aminophosphoramidates and aminoalkylphosphoramidates, phosphorodiamidates, thionophosphoramidates, alkylthiophosphates, alkylthiophosphotriester, selenophosphates and borophosphates, analogs with normal 3'-5' connection, 2'-5' connection and analogs with reverse polarity (wherein one or more internucleotide connections are 3' to 3', 5' to 5' or 2' to 2' connection).

[0221] Nucleic acids may include polynucleotide backbones formed by short-chain alkyl or cycloalkyl nucleoside bonds, mixed heteroatoms, and alkyl or cycloalkyl nucleoside bonds, or one or more short-chain heteroatomic or heterocyclic nucleoside bonds. These may include those with morpholino linkages (partially formed by the sugar portion of the nucleoside); siloxane backbones; sulfide, sulfoxide and sulfone backbones; formacetyl and thiocarbamate backbones; methylenecarbamate and thiocarbamate backbones; ribose acetyl backbones; olefin-containing backbones; aminosulfonate backbones; methyleneimino and methylenehydrazine backbones; sulfonate and sulfonamide backbones; amide backbones; and those with mixed N, O, S and CH 2 Others that are part of the components.

[0222] Nucleic acid can include nucleic acid mimics.Term " mimics " can be intended to include polynucleotides in which only furanose ring or both furanose ring and internucleotide bond are replaced by non-furanose groups, and only the replacement of furanose ring can also be called sugar substitute (surrogate).Heterocyclic base part or modified heterocyclic base part can be maintained to hybridize with appropriate target nucleic acid.A kind of such nucleic acid can be peptide nucleic acid (PNA).In PNA, the sugar backbone of polynucleotide can be replaced by amide-containing backbone, particularly replaced by aminoethylglycine backbone.Nucleotide can be retained and directly or indirectly combined with the nitrogen-nitrogen atom of the amide part of backbone.The backbone in PNA compound can include two or more aminoethylglycine units connected, which makes PNA have the backbone containing amide.Heterocyclic base part can be directly or indirectly combined with the nitrogen-nitrogen atom of the amide part of backbone.

[0223] Nucleic acid can include morpholino backbone structure.For example, nucleic acid can include 6-membered morpholino rings replacing ribose rings.In some of these embodiments, phosphorodiamidate or other non-phosphodiester internucleoside bonds can replace phosphodiester bonds.

[0224] Nucleic acid can include morpholino units (e.g., morpholino nucleic acids) having a connection to a heterocyclic base attached to a morpholino ring. A linking group can connect the morpholino monomer units in morpholino nucleic acids. Oligomeric compounds based on nonionic morpholinos can have less undesirable interactions with cell proteins. Polynucleotides based on morpholinos can be nonionic nucleic acid mimics. Various compounds within the morpholino category can be connected using different linking groups. Polynucleotide mimics of other categories can refer to cyclohexenyl nucleic acids (CeNA). The furanose rings commonly present in nucleic acid molecules can be replaced by cyclohexenyl rings. Phosphoramidite monomers protected by CeNA DMT can be prepared using phosphoramidite chemistry and are used for oligomeric compound synthesis. CeNA monomers are incorporated into nucleic acid chains to increase the stability of DNA / RNA hybrids. CeNA oligoadenylates can form complexes with nucleic acid complements, with stability similar to natural complexes. Additional modifications may include locked nucleic acids (LNA) in which the 2'-hydroxyl group is linked to the 4' carbon atom of the sugar ring, thereby forming a 2'-C,4'-C-oxymethylene linkage, thereby forming a bicyclic sugar moiety. The linkage may be a methylene (-CH 2 -), a group bridging the 2' oxygen atom and the 4' carbon atom, wherein n is 1 or 2. LNA and LNA analogs can show very high duplex thermal stability (Tm = +3 to +10°C) with complementary nucleic acids, stability to 3'-exonuclease degradation and good solubility.

[0225] Nucleic acids can also include modifications or substitutions of nucleic acid bases (commonly referred to as "bases"). As used herein, "unmodified" or "natural" nucleic acid bases can include purine bases (e.g., adenine (A) and guanine (G)), and pyrimidine bases (e.g., thymine (T), cytosine (C) and uracil (U)). Modified nucleic acid bases can include other synthetic and natural nucleic acid bases, such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C≡C-CH3) uracil and cytosine and other alkynyl derivatives of pyrimidine bases, 6-azo Uracil, cytosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halogen, 8-amino, 8-thio, 8-thioalkyl, 8-hydroxy and other 8-substituted adenines and guanines, 5-halogen, especially 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-amino-adenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine and 3-deazaguanine and 3-deazaadenine. The modified nucleic acid bases may include tricyclic pyrimidines such as phenoxazine cytidine (1H-pyrimido (5,4-b) (1,4) benzoxazin-2 (3H) -one), phenothiazine cytidine (1H-pyrimido (5,4-b) (1,4) benzothiazin-2 (3H) -one), G-clamps such as substituted phenoxazine cytidine (e.g., 9-(2-aminoethoxy)-H-pyrimido (5,4-(b) (1,4) benzoxazin-2 (3H) -one), phenoxazine cytidine (1H-pyrimido (5,4-b) (1,4) benzothiazin-2 (3H) -one), Thiazide cytidine (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), G-clamps such as substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido(4,5-b)indol-2-one), pyridoindole cytidine (H-pyrido(3',2':4,5)pyrrolo[2,3-d]pyrimidin-2-one).

[0226] As used herein, the term "sample" may refer to a composition comprising a target. Suitable samples for analysis by the disclosed methods, devices and systems include cells, tissues, organs or organisms. In some embodiments, the sample includes an original or untreated sample, such as a whole cell, a whole cell colony or a whole tissue. In some embodiments, the sample includes a separated cell or cell extract, or a fraction thereof containing nucleic acid, such as a separated nucleic acid, or a composition comprising an enriched or separated nucleic acid. In some embodiments, the sample includes a fixed tissue, a cell or a fraction thereof containing nucleic acid. In some embodiments, the sample includes a frozen tissue, a cell or a fraction thereof containing nucleic acid. In some embodiments, the sample includes a solution containing nucleic acid. In some embodiments, the sample includes a solution containing nucleic acid. In some embodiments, the sample includes nucleic acid in solid form, such as freeze-dried nucleic acid, etc. In some embodiments, the sample identifier sequence is a sequence that can provide information related to the sample to the user (such as identifying a specific sample or distinguishing a sample from another sample). The sample identifier sequence can be a unique nucleic acid sequence of 20 base pairs or less, such as 20, 18, 15, 12, 10, 9, 8, 7, 6, 5, 4 or 3 base pairs, from a unique set of distinct sequences.

[0227] As used herein, the term "sampling device" or "device" may refer to a device that can take a portion of a sample and / or place the portion on a substrate. The sampling device may refer to, for example, a fluorescence activated cell sorter (FACS) machine, a cell sorter, a biopsy needle, a biopsy device, a tissue sectioning device, a microfluidic device, a blade grid, and / or an ultramicrotome.

[0228] As used herein, the term "solid support" may refer to a discrete solid or semi-solid surface to which more than one barcode (e.g., a random barcode) may be attached. A solid support may include any type of solid, porous or hollow sphere, ball, bearing, cylinder or other similar configuration, including plastic, ceramic, metal or polymeric material (e.g., hydrogel), on which nucleic acids may be fixed (e.g., covalently or non-covalently). A solid support may include discrete particles that may be spherical (e.g., microspheres) or have a non-spherical or irregular shape, such as a cubic, rectangular, conical, cylindrical, conical, elliptical or disc-shaped, etc. The shape of a bead may be non-spherical. More than one solid support spaced apart in an array may not include a substrate. A solid support may be used interchangeably with the term "bead".

[0229] As used herein, the term "random barcode" may refer to a polynucleotide sequence comprising a tag of the present disclosure. A random barcode may be a polynucleotide sequence that can be used for random barcoding. A random barcode may be used to quantify a target in a sample. A random barcode may be used to control errors that may occur after the tag is associated with the target. For example, a random barcode may be used to assess amplification or sequencing errors. A random barcode associated with a target may be referred to as a random barcode-target or a random barcode-tag-target.

[0230] As used herein, the term "gene-specific random barcode" may refer to a polynucleotide sequence comprising a tag and a gene-specific target binding region. A random barcode may be a polynucleotide sequence that can be used for random barcoding. A random barcode may be used to quantify a target in a sample. A random barcode may be used to control errors that may occur after the tag is associated with the target. For example, a random barcode may be used to assess amplification or sequencing errors. A random barcode associated with a target may be referred to as a random barcode-target or a random barcode-tag-target.

[0231] As used herein, the term "stochastic barcoding" can refer to random labeling (e.g., barcoding) of nucleic acids. Stochastic barcoding can utilize a recursive Poisson strategy to associate and quantify labels associated with a target. As used herein, the term "stochastic barcoding" can be used interchangeably with "random labeling."

[0232] As used herein, the term "target" may refer to a composition that can be associated with a barcode (e.g., a random barcode). Exemplary suitable targets for analysis by the disclosed methods, devices, and systems include oligonucleotides, DNA, RNA, mRNA, microRNA, tRNA, etc. The target may be single-stranded or double-stranded. In some embodiments, the target may be a protein, peptide, or polypeptide. In some embodiments, the target is a lipid. As used herein, "target" may be used interchangeably with "species".

[0233] As used herein, the term "reverse transcriptase" refers to an enzyme with reverse transcriptase activity (i.e., catalyzing the synthesis of DNA from an RNA template). Typically, such enzymes include, but are not limited to, retroviral reverse transcriptases, retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retrotranscriptases, bacterial reverse transcriptases, type II intron-derived reverse transcriptases, and their mutants, variants, or derivatives. Non-retroviral reverse transcriptases include non-LTR retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retrotranscriptases, and type II intron reverse transcriptases. Examples of type II intron reverse transcriptases include Lactococcus lactis LI.LtrB intron reverse transcriptase, Thermosynechococcus elongatus TeI4c intron reverse transcriptase, or Geobacillus stearothermophilus GsI-IIC intron reverse transcriptase. Other classes of reverse transcriptases can include many types of non-retroviral reverse transcriptases (ie, retrovirals, group II introns, and diversity-generating retroelements, among others).

[0234] The terms "universal adapter primer," "universal primer adapter," or "universal adapter sequence" are used interchangeably to refer to a nucleotide sequence that can be used to hybridize with a barcode (e.g., a random barcode) to generate a gene-specific barcode. The universal adapter sequence can be, for example, a known sequence that is common to all barcodes used in the methods of the present disclosure. For example, when more than one target is labeled using the methods disclosed herein, each target-specific sequence can be connected to the same universal adapter sequence. In some embodiments, more than one universal adapter sequence can be used in the methods disclosed herein. For example, when more than one target is labeled using the methods disclosed herein, at least two target-specific sequences are connected to different universal adapter sequences. The universal adapter primer and its complementary sequence can be included in two oligonucleotides, one of which contains a target-specific sequence and the other oligonucleotide contains a barcode. For example, the universal adapter sequence can be a portion of an oligonucleotide containing a target-specific sequence to generate a nucleotide sequence complementary to a target nucleic acid. A second oligonucleotide containing a barcode and a complementary sequence to the universal adapter sequence can hybridize with the nucleotide sequence and generate a target-specific barcode (e.g., a target-specific random barcode). In some embodiments, the universal adapter primer has a different sequence than the universal PCR primer used in the methods of the disclosure.

[0235] As used herein, the term "cell" has the common and customary meanings understood by those of ordinary skill in the art according to this specification. It may refer to one or more cells. In some embodiments, the cell is a normal cell, for example, a human cell at different stages of development, or a human cell from different organs or tissue types. In some embodiments, the cell is a non-human cell, for example, other types of mammalian cells (e.g., mice, rats, pigs, dogs, cows and horses). In some embodiments, the cell is other types of animal or plant cells. In other embodiments, the cell may be any prokaryotic or eukaryotic cell. Permeabilized cells are cells comprising an opening in the cell membrane and a nuclease of sufficient size, which allows the protein binding reagent and the fusion protein (respectively, or optionally, associated with each other) to diffuse through the opening to enter the nucleus.

[0236] Barcode

[0237] Barcoding, such as random barcoding, has been described in, for example, Fu et al., Proc Natl Acad Sci U.S.A., 2011 May 31, 108(22):9026-31; U.S. Patent Application Publication No. US2011 / 0160078; Fan et al., Science, 2015 February 6, 347(6222):1258367; U.S. Patent Application Publication No. US2015 / 0299784; and PCT Application Publication No. WO2015 / 031691; the contents of each of these, including any supporting or supplementary information or materials, are incorporated herein by reference in their entirety. In some embodiments, the barcodes disclosed herein can be random barcodes, which can be polynucleotide sequences that can be used to randomly label (e.g., barcodes, tags) a target. A barcode can be referred to as a stochastic barcode if the ratio of the number of different barcode sequences of a stochastic barcode to the number of occurrences of any target to be labeled can be or can be about 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or a number or range between any two of these values. A target can be an mRNA species including mRNA molecules having the same or nearly the same sequence. A barcode can be referred to as a stochastic barcode if the ratio of the number of different barcode sequences of a stochastic barcode to the number of occurrences of any target to be labeled is at least or at most 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100: 1. The barcode sequences of a stochastic barcode can be referred to as molecular markers.

[0238] The barcode (e.g., random barcode) may include one or more labels. Exemplary labels may include universal labels, cell labels, barcode sequences (e.g., molecular labels), sample labels, plate labels, spatial labels, and / or pre-spatial labels. Fig.11An exemplary barcode 1104 with spatial markers is shown. The barcode 1104 may include a 5' amine that allows the barcode to be attached to a bead 1108. The barcode may contain universal markers, dimensional markers, spatial markers, cellular markers, and / or molecular markers. The order of the different markers (including but not limited to universal markers, dimensional markers, spatial markers, cellular markers, and molecular markers) in the barcode may vary. For example, Fig.11 As shown in , the universal label can be the label of the most 5' side (5'-most label), and the molecular label can be the label of the most 3' side (3'-most label). Spatial label, dimensional label and cell label can be in any order. In some embodiments, universal label, spatial label, dimensional label, cell label and molecular label are in any order. Barcode can include target binding region. Target binding region can interact with the target (for example, target nucleic acid, RNA, mRNA, DNA) in sample. For example, the target binding region can include oligo (dT) sequence that can interact with the poly (A) tail of mRNA. In some cases, the label of barcode (for example, universal label, dimensional label, spatial label, cell label and barcode sequence) can be separated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 or more nucleotides.

[0239] Mark (for example, cell marker) can comprise a group of unique defined length nucleic acid subsequences, for example, every seven nucleotides (equivalent to the number of bits used in some Hamming error correction codes), which can be designed to provide error correction capabilities. The error correction subsequence group comprising seven nucleotide sequences can be designed so that any paired combination of the sequences in the group exhibits a defined "genetic distance" (or mismatched base number), for example, a group of error correction subsequences can be designed to exhibit a genetic distance of three nucleotides. In this case, the examination of the error correction sequence in the sequence data group of the target nucleic acid molecule of the mark (described in more detail below) can allow people to detect or correct amplification errors or sequencing errors. In some embodiments, the length of the nucleic acid subsequence for generating the error correction code can vary, for example, their length can be following or can be about following: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 31, 40, 50 nucleotides or the nucleotides of the number or range between any two of these values. In some embodiments, nucleic acid subsequences of other lengths can be used to generate error correction codes.

[0240] The barcode may include a target binding region. The target binding region may interact with a target in a sample. The target may be or include: ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNA each containing a poly (A) tail, or any combination thereof. In some embodiments, more than one target may include deoxyribonucleic acid (DNA).

[0241] In some embodiments, the target binding region may include an oligo (dT) sequence that can interact with the poly (A) tail of the mRNA. One or more markers of the barcode (e.g., universal markers, dimensional markers, spatial markers, cell markers, and barcode sequences (e.g., molecular markers)) may be separated from another or two remaining markers of the barcode by a spacer. The spacer may be, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 or more nucleotides. In some embodiments, no marker is separated by a spacer in the marker of the barcode.

[0242] General Tags

[0243] Barcode can comprise one or more universal markers.In some embodiments, one or more universal markers can be the same for all barcodes in the barcode group attached to a given solid support.In some embodiments, one or more universal markers can be the same for all barcodes attached to more than one bead.In some embodiments, universal marker can include a nucleic acid sequence that can be hybridized with a sequencing primer.Sequencing primer can be used to sequence the barcode comprising the universal marker.Sequencing primer (for example, universal sequencing primer) can include sequencing primers related to high-throughput sequencing platform.In some embodiments, universal marker can include a nucleic acid sequence that can be hybridized with a PCR primer.In some embodiments, universal marker can include a nucleic acid sequence that can be hybridized with a sequencing primer and a PCR primer.The nucleic acid sequence of the universal marker that can be hybridized with a sequencing primer or a PCR primer can be referred to as a primer binding site.The universal marker can include a sequence that can be used to initiate the transcription of a barcode.The universal marker can include a sequence that can be used to extend a barcode or a region within a barcode. The length of the universal tag can be or can be about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides or a number or range of nucleotides between any two of these values. For example, the universal tag can include at least about 10 nucleotides. The length of the universal tag can be at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200 or 300 nucleotides. In some embodiments, a cleavable linker or modified nucleotide can be part of the universal tag sequence to enable the barcode to be cleaved from the support.

[0244] Dimension tagging

[0245] The barcode may include one or more dimensional markers. In some embodiments, the dimensional marker may include a nucleic acid sequence that provides information about the dimension in which the marker (e.g., random marker) occurs. For example, the dimensional marker may provide information about the time at which the target is barcoded. The dimensional marker may be associated with the time of barcoding (e.g., random barcoding) in the sample. The dimensional marker may be activated at the time of the marker. Different dimensional markers may be activated at different times. The dimensional marker provides information about the order in which the target, the target group, and / or the sample are barcoded. For example, a population of cells may be barcoded in the G0 phase of the cell cycle. In the G1 phase of the cell cycle, the cell may be pulsed again with a barcode (e.g., a random barcode). In the S phase of the cell cycle, the cell may be pulsed again with a barcode, and so on. The barcode of each pulse (e.g., each period of the cell cycle) may include different dimensional markers. In this way, the dimensional marker provides information about which targets are marked in which period of the cell cycle. Dimensional markers can interrogate many different biological stages. Exemplary biological events may include, but are not limited to, cell cycle, transcription (e.g., transcription initiation), and transcript degradation. In another example, a sample (e.g., a cell, a population of cells) may be labeled before and / or after treatment with a drug and / or therapy. Changes in the copy number of different targets may indicate the response of a sample to a drug and / or therapy.

[0246] Dimensional markers can be activatable. Activatable dimensional markers can be activated at a specific time point. Activatable markers can be, for example, constitutively activated (e.g., not closed). Activatable dimensional markers can be, for example, reversibly activated (e.g., activatable dimensional markers can be turned on and off). Dimensional markers can be reversibly activated, for example, at least 1 time, 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times or more. Dimensional markers can be reversibly activated, for example, at least 1 time, 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times or more. In some embodiments, dimensional markers can be activated by fluorescence, light, chemical events (e.g., cleavage, connecting another molecule, adding modifications (e.g., pegylation, sumoylate, acetylation, methylation, deacetylation, demethylation), photochemical events (e.g., light imprisonment (photocaging)) and the introduction of non-natural nucleotides.

[0247] In some embodiments, the dimension mark can be the same for all bar codes (e.g., random bar codes) attached to a given solid support (e.g., beads), but different for different solid supports (e.g., beads). In some embodiments, at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99% or 100% of the bar codes on the same solid support can include the same dimension mark. In some embodiments, at least 60% of the bar codes on the same solid support can include the same dimension mark. In some embodiments, at least 95% of the bar codes on the same solid support can include the same dimension mark.

[0248] More than one solid support (e.g., beads) may be present with up to 10 6 The length of the dimension marker can be as follows or can be about as follows: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides or a number or range of nucleotides between any two of these values. The length of the dimension marker can be at least as follows or at most as follows: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200 or 300 nucleotides. The dimension marker can include about 5 to about 200 nucleotides. The dimension marker can include about 10 to about 150 nucleotides. The dimension marker can include a length of about 20 to about 125 nucleotides.

[0249] Space Marking

[0250] The barcode may include one or more spatial markers. In some embodiments, the spatial marker may include a nucleic acid sequence that provides information about the spatial orientation of the target molecule associated with the barcode. The spatial marker may be associated with a coordinate in the sample. The coordinate may be a fixed coordinate. For example, the coordinate may be fixed relative to a substrate. The spatial marker may refer to a two-dimensional or three-dimensional grid. The coordinate may be fixed relative to a landmark. The landmark may be identified in space. The landmark may be a structure that may be imaged. The landmark may be a biological structure, such as an anatomical landmark. The landmark may be a cell landmark, such as an organelle. The landmark may be a non-natural landmark, such as a structure with an identifiable identifier (such as a color code, a barcode, a magnetic property, fluorescence, radioactivity, or a unique size or shape). The spatial marker may be associated with a physical partition (e.g., a hole, a container, or a droplet). In some embodiments, more than one spatial marker may be used together to encode one or more positions in space.

[0251] The spatial tag can be the same for all barcodes attached to a given solid support (e.g., beads), but different for different solid supports (e.g., beads). In some embodiments, the percentage of barcodes on the same solid support containing the same spatial tag can be the following or can be about the following: 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100% or a number or range between any two of these values. In some embodiments, the percentage of barcodes on the same solid support containing the same spatial tag can be at least or at most 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99% or 100%. In some embodiments, at least 60% of the barcodes on the same solid support can contain the same spatial tag. In some embodiments, at least 95% of the barcodes on the same solid support can contain the same spatial tag.

[0252] More than one solid support (e.g., beads) may be present with up to 10 6 The length of spatial marker can be as follows or can be about as follows: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides or the nucleotides of the number or scope between any two of these values. The length of spatial marker can be at least as follows or at most as follows: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200 or 300 nucleotides. Spatial marker can comprise about 5 to about 200 nucleotides. Spatial marker can comprise about 10 to about 150 nucleotides. Spatial marker can comprise a length of about 20 to about 125 nucleotides.

[0253] Cell labeling

[0254] Barcodes (e.g., random barcodes) may include one or more cell markers. In some embodiments, cell markers may include nucleic acid sequences that provide information for determining which target nucleic acid is derived from which cell. In some embodiments, cell markers are identical for all barcodes attached to a given solid support (e.g., beads), but are different for different solid supports (e.g., beads). In some embodiments, the percentage of barcodes on the same solid support comprising the same cell markers may be as follows or may be approximately as follows: 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100% or a number or range between any two of these values. In some embodiments, the percentage of barcodes on the same solid support comprising the same cell markers may be as follows or may be approximately as follows: 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99% or 100%. For example, at least 60% of the barcodes on the same solid support may include the same cell marker. As another example, at least 95% of the barcodes on the same solid support can comprise the same cellular marker.

[0255] More than one solid support (e.g., beads) may be present with up to 10 6 The length of the cell marker can be as follows or can be about as follows: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides or the nucleotides of the number or range between any two of these values. The length of the cell marker can be at least as follows or at most as follows: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200 or 300 nucleotides. For example, the cell marker can include about 5 to about 200 nucleotides. As another example, the cell marker can include about 10 to about 150 nucleotides. As another example, the cell marker can include a length of about 20 to about 125 nucleotides.

[0256] Barcode sequence

[0257] The barcode can comprise one or more barcode sequences. In some embodiments, the barcode sequence can comprise a nucleic acid sequence that provides identification information for a particular type of target nucleic acid species that hybridizes to the barcode. The barcode sequence can comprise a nucleic acid sequence that provides a counter (e.g., provides a rough approximation) for a particular occurrence of a target nucleic acid species that hybridizes to a barcode (e.g., a target binding region).

[0258] In some embodiments, a set of distinct barcode sequences are attached to a given solid support (e.g., a bead). In some embodiments, there may be or about the following: 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 unique molecular marker sequences or a number or range of unique molecular marker sequences between any two of these values. For example, more than one barcode can include about 6561 barcode sequences with different sequences. As another example, more than one barcode can include about 65536 barcode sequences with different sequences. In some embodiments, there can be at least the following or at most the following: 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 or 10 9 A unique barcode sequence. The unique molecular marker sequence can be attached to a given solid support (e.g., a bead). In some embodiments, the unique molecular marker sequence is partially or fully contained by a particle (e.g., a hydrogel bead).

[0259] In different embodiments, the length of the barcode can be different. For example, the length of the barcode can be or can be about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides or a number or range of nucleotides between any two of these values. As another example, the length of the barcode can be at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides.

[0260] Molecular markers

[0261] The barcode (e.g., a random barcode) can include one or more molecular markers. The molecular marker can include a barcode sequence. In some embodiments, the molecular marker can include a nucleic acid sequence that provides identification information for a specific type of target nucleic acid species that hybridizes to the barcode. The molecular marker can include a nucleic acid sequence that provides a counter for a specific occurrence of a target nucleic acid species that hybridizes to a barcode (e.g., a target binding region).

[0262] In some embodiments, a set of distinct molecular markers are attached to a given solid support (e.g., a bead). In some embodiments, there may be or are about 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 or a number or range of unique molecular marker sequences between any two of these values. For example, more than one barcode may include about 6561 molecular markers with different sequences. As another example, more than one barcode may include about 65536 molecular markers with different sequences. In some embodiments, there may be at least or at most 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 or 10 9 A barcode with a unique molecular marker sequence can be attached to a given solid support (e.g., a bead).

[0263] For barcoding using more than one random barcode (e.g., random barcoding), the ratio of the number of different molecular marker sequences to the number of occurrences of any target can be or can be about 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or a number or range between any two of these values. The target can be an mRNA species including mRNA molecules having the same or nearly the same sequence. In some embodiments, the ratio of the number of different molecular marker sequences to the number of occurrences of any target is at least or at most 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100:1.

[0264] The length of a molecular marker can be or can be about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides or a number or range of nucleotides between any two of these values. The length of a molecular marker can be at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200 or 300 nucleotides.

[0265] Target binding region

[0266] The barcode may include one or more target binding regions, such as capture probes. In some embodiments, the target binding region may hybridize with a target of interest. In some embodiments, the target binding region may include a nucleic acid sequence that specifically hybridizes with a target (e.g., a target nucleic acid, a target molecule, such as a cell nucleic acid to be analyzed) (e.g., specifically hybridizes with a specific gene sequence). In some embodiments, the target binding region may include a nucleic acid sequence that can be attached (e.g., hybridized) to a specific position of a specific target nucleic acid. In some embodiments, the target binding region may include a nucleic acid sequence that can specifically hybridize with a restriction enzyme site overhang (e.g., an EcoRI sticky end overhang). The barcode may then be connected to any nucleic acid molecule comprising a sequence complementary to a restriction site overhang.

[0267] In some embodiments, the target binding region may include a non-specific target nucleic acid sequence. Non-specific target nucleic acid sequences may refer to sequences that can bind to more than one target nucleic acid independently of the specific sequence of the target nucleic acid. For example, the target binding region may include a random polymer sequence, a poly (dA) sequence, a poly (dT) sequence, a poly (dG) sequence, a poly (dC) sequence, or a combination thereof. For example, the target binding region may be an oligo (dT) sequence hybridized with a poly (A) tail on an mRNA molecule. The random polymer sequence may be, for example, a random dimer, trimer, tetramer, pentamer, hexamer, heptamer, octamer, nonamer, decamer, or a higher polymer sequence of any length. In some embodiments, for all barcodes attached to a given bead, the target binding region is the same. In some embodiments, for more than one barcode attached to a given bead, the target binding region may include two or more different target binding sequences. The length of the target binding region can be or can be about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides, or a number or range of nucleotides between any two of these values. The length of the target binding region can be up to about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 or more nucleotides. For example, a reverse transcriptase (such as Moloney murine leukemia virus (MMLV) reverse transcriptase) can be used to reverse transcribe an mRNA molecule to produce a cDNA molecule with a poly (dC) tail. The barcode can include a target binding region with a poly (dG) tail. After base pairing between the poly (dG) tail of the barcode and the poly (dC) tail of the cDNA molecule, the reverse transcriptase converts the template strand from the cellular RNA molecule to the barcode and continues to replicate to the 5' end of the barcode. By doing so, the resulting cDNA molecule contains the sequence of the barcode (such as a molecular tag) on ​​the 3' end of the cDNA molecule.

[0268] In some embodiments, the target binding region can include oligo (dT), which can hybridize with mRNA comprising polyadenylated ends. The target binding region can be gene specific. For example, the target binding region can be configured to hybridize with a specific region of the target. The length of the target binding region can be as follows or can be about as follows: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 27, 28, 29, 30 nucleotides or a number or range of nucleotides between any two of these values. The length of the target binding region can be at least or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 27, 28, 29, or 30 nucleotides. The length of the target binding region can be about 5-30 nucleotides. When the barcode comprises a gene-specific target binding region, the barcode may be referred to herein as a gene-specific barcode.

[0269] Orientation Property

[0270] Random barcodes (e.g., random barcodes) can include one or more directional properties that can be used to orient (e.g., align) the barcodes. The barcodes can include portions for isoelectric focusing. Different barcodes can include different isoelectric focusing points. When these barcodes are introduced into a sample, the sample can undergo isoelectric focusing to orient the barcodes in a known manner. In this way, the directional properties can be used to develop known mappings of barcodes in a sample. Exemplary directional properties can include electrophoretic mobility (e.g., based on the size of the barcode), isoelectric point, spin, conductivity, and / or self-assembly. For example, a barcode with a directional property of self-assembly can self-assemble into a specific orientation (e.g., a nucleic acid nanostructure) when activated.

[0271] AffinityProperty

[0272] Barcodes (e.g., random barcodes) can include one or more affinity properties. For example, spatial labels can include affinity properties. Affinity properties can include chemical and / or biological parts that can promote the binding of barcodes to another entity (e.g., cell receptors). For example, affinity properties can include antibodies, for example, antibodies specific to a specific part (e.g., receptor) on a sample. In some embodiments, antibodies can guide barcodes to specific cell types or molecules. Targets at and / or near specific cell types or molecules can be marked (e.g., randomly marked). In some embodiments, affinity properties can provide spatial information beyond the nucleotide sequence of the spatial label because antibodies can guide barcodes to specific locations. The antibody can be a therapeutic antibody, such as a monoclonal antibody or a polyclonal antibody. The antibody can be humanized or chimeric. The antibody can be a naked antibody or a fusion antibody.

[0273] Antibodies can be full-length (ie, naturally occurring or formed by normal immunoglobulin gene fragment recombination processes) immunoglobulin molecules (eg, IgG antibodies) or immunologically active (ie, specifically binding) portions of immunoglobulin molecules (such as antibody fragments).

[0274] The antibody fragment can be, for example, a part of an antibody, such as F(ab')2, Fab', Fab, Fv, sFv, etc. In some embodiments, the antibody fragment can bind to the same antigen recognized by the full-length antibody. The antibody fragment can include a separated fragment consisting of the variable region of the antibody, such as a "Fv" fragment consisting of the variable region of the heavy chain and the light chain, and a recombinant single-chain polypeptide molecule ("scFv protein") in which the light chain and the heavy chain variable region are connected by a peptide linker. Exemplary antibodies can include, but are not limited to, cancer cell antibodies, virus antibodies, antibodies that bind to cell surface receptors (CD8, CD34, CD45), and therapeutic antibodies.

[0275] Universal adapter primer

[0276] The barcode can include one or more universal adapter primers. For example, a gene-specific barcode (such as a gene-specific random barcode) can include a universal adapter primer. A universal adapter primer can refer to a universal nucleotide sequence across all barcodes. A universal adapter primer can be used to construct a gene-specific barcode. The length of the universal adapter primer can be as follows or can be about as follows: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 27, 28, 29, 30 nucleotides or a number or range of nucleotides between any two of these values. The length of the universal adapter primer can be at least or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 27, 28, 29, or 30 nucleotides. The length of the universal adapter primer can be 5-30 nucleotides.

[0277] Connectors

[0278] When the barcode comprises more than one type of marker (e.g., more than one cell marker or more than one barcode sequence, such as a molecular marker), the markers may be interspersed with adapter marker sequences. The length of the adapter marker sequence may be at least about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 or more nucleotides. The length of the adapter marker sequence may be at most about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 or more nucleotides. In some cases, the length of the adapter marker sequence is 12 nucleotides. The adapter marker sequence can be used to facilitate the synthesis of the barcode. The adapter marker may include an error correction (e.g., Hamming) code.

[0279] Solid support

[0280] In some embodiments, the barcodes disclosed herein (such as random barcodes) can be associated with a solid support. The solid support can be, for example, a synthetic particle. In some embodiments, some or all barcode sequences (such as molecular markers of random barcodes (e.g., first barcode sequences)) of more than one barcode on a solid support (e.g., more than one barcode) differ by at least one nucleotide. The cell markers of the barcodes on the same solid support can be the same. The cell markers of the barcodes on different solid supports can differ by at least one nucleotide. For example, the first cell marker of the first more than one barcode on the first solid support can have the same sequence, and the second cell marker of the second more than one barcode on the second solid support can have the same sequence. The first cell marker of the first more than one barcode on the first solid support and the second cell marker of the second more than one barcode on the second solid support can differ by at least one nucleotide. The cell marker can be, for example, about 5-20 nucleotides long. The barcode sequence can be, for example, about 5-20 nucleotides long. The synthetic particle can be, for example, a bead.

[0281] The beads can be, for example, silica beads, controlled pore glass beads, magnetic beads, Dynabeads, Sephadex / agarose beads, cellulose beads, polystyrene beads, or any combination thereof. The beads can include materials such as polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic substances, ceramics, plastics, glass, methylstyrene, acrylic polymers, titanium, latex, agarose gel, cellulose, nylon, silicone, or any combination thereof.

[0282] In some embodiments, the beads can be polymer beads (e.g., deformable beads or gel beads) functionalized with barcodes or random barcodes (such as gel beads from 10X Genomics (San Francisco, CA)). In some embodiments, the gel beads can include a polymer-based gel. The gel beads can be produced, for example, by encapsulating one or more polymer precursors into droplets. The gel beads can be produced after exposing the polymer precursors to an accelerator (e.g., tetramethylethylenediamine (TEMED)).

[0283] In some embodiments, the particles can be destructible (e.g., soluble, degradable). For example, the polymer beads can dissolve, melt or degrade, for example, under desired conditions. The desired conditions can include environmental conditions. The desired conditions can cause the polymer beads to dissolve, melt or degrade in a controlled manner. The gel beads can dissolve, melt or degrade due to chemical stimulation, physical stimulation, biological stimulation, thermal stimulation, magnetic stimulation, electrical stimulation, light stimulation or any combination thereof.

[0284] For example, analyte and / or reagent (such as oligonucleotide barcode) can be coupled / fixed to the inner surface of gel beads (for example, via the diffusion of oligonucleotide barcode and / or the material for producing oligonucleotide barcode) and / or the outer surface of gel beads or any other microcapsule described herein. Coupling / fixation can be via any form of chemical bonding (for example, covalent bond, ionic bond) or physical phenomenon (for example, van der Waals force, dipole-dipole interaction, etc.). In some embodiments, the coupling / fixation of reagents described herein to gel beads or any other microcapsule can be reversible, such as, for example, via unstable part (for example, via chemical crosslinking agent, including chemical crosslinking agent described herein). After applying stimulation, the unstable part can be cleaved and release the fixed reagent. In some embodiments, the unstable part is a disulfide bond. For example, in the case where the oligonucleotide barcode is fixed to the gel beads via a disulfide bond, exposing the disulfide bond to a reducing agent can cleave the disulfide bond and release the oligonucleotide barcode from the beads. The labile moiety can be included as part of a gel bead or microcapsule, as part of a chemical linker that connects a reagent or analyte to a gel bead or microcapsule, and / or as part of a reagent or analyte. In some embodiments, at least one barcode of more than one barcode can be immobilized on a particle, partially immobilized on a particle, encapsulated in a particle, partially encapsulated in a particle, or any combination thereof.

[0285] In some embodiments, the gel beads may include a wide range of different polymers, including but not limited to: polymers, thermosensitive polymers, photosensitive polymers, magnetic polymers, pH sensitive polymers, salt sensitive polymers, chemical sensitive polymers, polyelectrolytes, polysaccharides, peptides, proteins and / or plastics. The polymer may include, but is not limited to, materials such as poly(N-isopropylacrylamide) (PNIPAAm), poly(styrene sulfonate) (PSS), poly(allylamine) (PAAm), poly(acrylic acid) (PAA), poly(ethyleneimine) (PEI), poly(bisallyldimethyl-ammonium chloride) (PDADMAC), poly(pyrrole) (poly(pyrolle), PPy), poly(vinylpyrrolidone) (PVPON), poly(vinylpyridine) (PVP), poly(methacrylic acid) (PMAA), poly(methyl methacrylate) (PMMA), polystyrene (PS), poly(tetrahydrofuran) (PTHF), poly(o-phthalaldehyde) (PPA), poly(hexylviologen) (PHV), poly(L-lysine) (PLL), poly(L-arginine) (PARG), and poly(lactic-co-glycolic acid) (PLGA).

[0286] Many chemical stimuli can be used to trigger the destruction, dissolution or degradation of beads. Examples of these chemical changes can include, but are not limited to, pH-mediated changes in the bead wall, chemical cleavage of the bead wall via crosslink bonds, triggered depolymerization of the bead wall, and bead wall conversion reactions. Bulk changes can also be used to trigger the destruction of beads.

[0287] Bulk or physical alteration of microcapsules by various stimuli also provides many advantages in designing capsules to release agents. Bulk or physical alteration occurs on a macroscopic scale, where bead rupture is the result of mechanical-physical forces caused by the stimulus. These processes can include, but are not limited to, pressure-induced rupture, bead wall melting, or changes in the porosity of the bead wall.

[0288] Biostimulation can also be used to trigger the destruction, dissolution or degradation of beads. Generally, biological triggers are similar to chemical triggers, but many examples use biomolecules or molecules common in living systems, such as enzymes, peptides, sugars, fatty acids, nucleic acids, etc. For example, beads can include polymers with peptide crosslinks that are sensitive to cleavage by specific proteases. More specifically, one example can include microcapsules containing GFLGK peptide crosslinks. Upon addition of a biological trigger (such as the protease cathepsin B), the peptide crosslinks of the shell wall are cleaved and the contents of the beads are released. In other cases, the protease can be heat-activated. In another example, the beads include a shell wall that includes cellulose. The addition of chitosan hydrolase serves as a biotrigger for cleavage of cellulose bonds, depolymerization of the shell wall and release of its internal contents.

[0289] The beads can also be induced to release their contents upon application of a thermal stimulus. Changes in temperature can cause various changes in the beads. Changes in heat can cause the beads to melt, causing the bead wall to disintegrate. In other cases, heat can increase the internal pressure of components within the beads, causing the beads to rupture or explode. In still other cases, heat can cause the beads to transform into a shrunken, dehydrated state. Heat can also act on thermosensitive polymers within the bead wall, causing the beads to break down.

[0290] Including magnetic nanobeads in the bead wall of the microcapsule can allow for triggered rupture of the beads and directing the beads into an array. The device of the present disclosure may include magnetic beads for any purpose. In one example, Fe 3 O 4 The nanoparticles are incorporated into polyelectrolyte-containing beads, which trigger rupture in the presence of an oscillating magnetic field stimulus.

[0291] The beads can also be destroyed, dissolved or degraded as a result of electrical stimulation. Similar to the magnetic particles described in the previous section, the electrosensitive beads can allow for triggered rupture of the beads as well as other functions such as alignment in an electric field, conductivity or redox reactions. In one example, beads containing electrosensitive materials align in an electric field so that the release of internal reagents can be controlled. In other examples, the electric field can induce redox reactions within the bead wall itself, which can increase porosity.

[0292] Light stimulation can also be used to disrupt the beads. Many light triggers are possible and can include systems using various molecules such as nanoparticles and chromophores that can absorb photons of a specific wavelength range. For example, metal oxide coatings can be used as capsule triggers. 2 UV irradiation of polyelectrolyte capsules can lead to disintegration of the bead wall. In yet another example, photoswitchable materials (such as azobenzene groups) can be incorporated into the bead wall. Upon application of UV or visible light, chemicals such as these undergo reversible cis-to-trans isomerization upon absorption of photons. In this regard, the incorporation of a photon switch produces a bead wall that can disintegrate or become more porous upon application of a light trigger.

[0293] For example, in Fig.12 In the non-limiting example of barcoding (e.g., random barcoding) shown in FIG, after introducing cells (such as single cells) into more than one microwell of a microwell array at block 1208, beads may be introduced into more than one microwell of a microwell array at block 1212. Each microwell may contain one bead. The beads may contain more than one barcode. The barcode may contain a 5' amine region attached to the bead. The barcode may contain a universal tag, a barcode sequence (e.g., a molecular tag), a target binding region, or any combination thereof.

[0294] The barcodes disclosed herein can be associated (e.g., attached) with a solid support (e.g., a bead). The barcodes associated with the solid support can each comprise a barcode sequence selected from the following group, the group comprising at least 100 or 1000 barcode sequences with unique sequences. In some embodiments, different barcodes associated with the solid support can comprise barcodes with different sequences. In some embodiments, a certain percentage of the barcodes associated with the solid support comprises the same cell marker. For example, the percentage can be below or about below: 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100% or a number or range between any two of these values. As another example, the percentage can be at least below or at most below: 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99% or 100%. In some embodiments, the barcodes associated with the solid support can have the same cell marker. The barcodes associated with different solid supports can have different cell markers selected from the group consisting of at least 100 or 1000 cell markers having unique sequences.

[0295] The barcodes disclosed herein can be associated (e.g., attached) to a solid support (e.g., a bead). In some embodiments, more than one target in a sample can be barcoded with a solid support including more than one synthetic particle associated with more than one barcode. In some embodiments, a solid support can include more than one synthetic particle associated with more than one barcode. The spatial labeling of more than one barcode on different solid supports can differ by at least one nucleotide. The solid support can include more than one barcode, for example, in two or three dimensions. The synthetic particle can be a bead. The bead can be a silica bead, a controlled pore glass bead, a magnetic bead, a Dynabead, a Sephadex / agarose gel bead, a cellulose bead, a polystyrene bead, or any combination thereof. The solid support can include a polymer, a matrix, a hydrogel, a needle array device, an antibody, or any combination thereof. In some embodiments, the solid support can float freely. In some embodiments, the solid support can be embedded in a semi-solid or solid array. The barcode may not be associated with the solid support. The barcode can be a single nucleotide. The barcode may be associated with the substrate.

[0296] As used herein, the terms "tethered," "attached," and "fixed" can be used interchangeably and can refer to covalent or non-covalent means for attaching a barcode to a solid support. Any of a variety of different solid supports can be used as a solid support for attaching pre-synthesized barcodes or for in situ solid phase synthesis of barcodes.

[0297] In some embodiments, the solid support is a bead. The bead may include one or more types of solid, porous or hollow spheres, balls, sockets, cylinders or other similar configurations on which nucleic acids can be fixed (e.g., covalently or non-covalently). The bead may be composed of, for example, plastic, ceramic, metal, polymeric material or any combination thereof. The bead may be or include spherical (e.g., microspheres) or discrete particles having a non-spherical or irregular shape, such as a cube, rectangle, cone, cylinder, conical, elliptical or disc-shaped, etc. In some embodiments, the shape of the bead may be non-spherical.

[0298] The beads may include a variety of materials, including but not limited to paramagnetic materials (e.g., magnesium, molybdenum, lithium, and tantalum), superparamagnetic materials (e.g., ferrite (Fe 3 O 4 ; magnetite) nanoparticles), ferromagnetic materials (e.g., iron, nickel, cobalt, some alloys thereof, and some rare earth metal compounds), ceramics, plastics, glass, polystyrene, silica, methylstyrene, acrylic polymers, titanium, latex, agarose gel, agarose, hydrogels, polymers, cellulose, nylon, or any combination thereof.

[0299] In some embodiments, the beads (e.g., beads to which labels are attached) are hydrogel beads. In some embodiments, the beads include a hydrogel.

[0300] Some embodiments disclosed herein include one or more particles (e.g., beads). Each particle can contain more than one oligonucleotide (e.g., barcode). Each of more than one oligonucleotide can contain a barcode sequence (e.g., a molecular marker sequence), a cell marker, and a target binding region (e.g., an oligo (dT) sequence, a gene-specific sequence, a random polymer, or a combination thereof). The cell marker sequence of each of more than one oligonucleotide can be the same. The cell marker sequences of the oligonucleotides on different particles can be different, so that the oligonucleotides on different particles can be identified. In different embodiments, the number of different cell marker sequences can be different. In some embodiments, the number of cell marker sequences may be or may be about the following: 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 10 ... 6 , 10 7 , 10 8 , 10 9, a number or range between any two of these values ​​or more. In some embodiments, the number of cell marker sequences can be at least the following or at most the following: 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000 6 , 10 7 , 10 8 or 10 9 In some embodiments, no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 or more oligonucleotides having the same cellular sequence are present in more than one particle. In some embodiments, at most 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10% or more of the more than one particle comprising oligonucleotides having the same cellular sequence. In some embodiments, more than one particle does not have the same cell marker sequence in all of them.

[0301] More than one oligonucleotide on each particle can comprise different barcode sequences (e.g., molecular markers). In some embodiments, the number of barcode sequences can be or can be about the following: 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000 6 , 10 7 , 10 8 , 10 9or a number or range between any two of these values. In some embodiments, the number of barcode sequences can be at least the following or at most the following: 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000, 100000 6 , 10 7 , 10 8 or 10 9 For example, at least 100 of the more than one oligonucleotides comprise different barcode sequences. As another example, in a single particle, at least 100, 500, 1000, 5000, 10000, 15000, 20000, 50000, a number or range between any two of these values, or more of the more than one oligonucleotides comprise different barcode sequences. Some embodiments provide more than one particle comprising a barcode. In some embodiments, the ratio of the occurrence (or copies or number) of the target to be labeled and the different barcode sequences can be at least 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 1:30, 1:40, 1:50, 1:60, 1:70, 1:80, 1:90 or more. In some embodiments, each of the more than one oligonucleotide further comprises a sample label, a universal label, or both. The particle can be, for example, a nanoparticle or a microparticle.

[0302] The size of the beads can be different. For example, the diameter of the beads can range from 0.1 micron to 50 microns. In some embodiments, the diameter of the beads can be or can be about 0.1 micron, 0.5 micron, 1 micron, 2 microns, 3 microns, 4 microns, 5 microns, 6 microns, 7 microns, 8 microns, 9 microns, 10 microns, 20 microns, 30 microns, 40 microns, 50 microns, or a number or range between any two of these values.

[0303] The diameter of the bead can be related to the diameter of the hole of the substrate. In some embodiments, the diameter of the bead can be longer or shorter than the diameter of the hole or can be longer or shorter than the diameter of the hole by about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100% or a number or range between any two of these values. The diameter of the bead can be related to the diameter of the cell (e.g., a single cell captured by the hole of the substrate). In some embodiments, the diameter of the bead can be longer or shorter than the diameter of the hole by at least or at most 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100%. The diameter of the bead can be related to the diameter of the cell (e.g., a single cell captured by the hole of the substrate). In some embodiments, the diameter of the bead can be longer or shorter than the diameter of the cell or can be about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 250%, 300% longer or shorter than the diameter of the cell, or a number or range between any two of these values. In some embodiments, the diameter of the bead can be at least or at most 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 250%, or 300% longer or shorter than the diameter of the cell.

[0304] The beads can be attached to a substrate and / or embedded in a substrate. The beads can be attached to a gel, a hydrogel, a polymer and / or a matrix and / or embedded in a gel, a hydrogel, a polymer and / or a matrix. The spatial position of the beads in a substrate (e.g., a gel, a matrix, a support or a polymer) can be identified using a spatial marker present on a barcode on the beads, which can be used as a positional address.

[0305] Examples of beads may include, but are not limited to, streptavidin beads, agarose beads, magnetic beads, microbeads, antibody-conjugated beads (e.g., anti-immunoglobulin microbeads), protein A-conjugated beads, protein G-conjugated beads, protein A / G-conjugated beads, protein L-conjugated beads, oligo(dT)-conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, and BcMag TM Carboxyl-terminated magnetic beads.

[0306] The beads can be associated with (e.g., impregnated with) quantum dots or fluorescent dyes to make them fluoresce in one fluorescent optical channel or more than one optical channel. The beads can be associated with iron oxide or chromium oxide to make them paramagnetic or ferromagnetic. The beads can be identifiable. For example, a camera can be used to image the beads. The beads can have a detectable code associated with the beads. For example, the beads can contain a barcode. The beads can change size, for example due to swelling in an organic or inorganic solution. The beads can be hydrophobic. The beads can be hydrophilic. The beads can be biocompatible.

[0307] The solid support (e.g., bead) can be visualized. The solid support can include a visualization label (e.g., a fluorescent dye). The solid support (e.g., bead) can be etched with an identifier (e.g., a number). The identifier can be visualized by imaging the bead.

[0308] The solid support may include soluble, semi-soluble or insoluble materials. When the solid support includes a linker, a scaffold, a building block or other reactive moiety attached thereto, the solid support may be referred to as "functionalized", and when the solid support lacks such reactive moieties attached thereto, the solid support may be referred to as "non-functionalized". The solid support may be free in solution, such as in a microtiter well; in a flow-through form, such as in a column; or used as a dipstick.

[0309] Solid supports can include films, paper, plastics, coated surfaces, flat surfaces, glass, slides, chips or any combination thereof. Solid supports can take the form of resins, gels, microspheres or other geometric configurations. Solid supports can include silicon dioxide chips, micron particles, nanoparticles, plates, arrays, capillaries, flat supports such as glass fiber filters, glass surfaces, metal surfaces (steel, gold, silver, aluminum, silicon and copper), glass supports, plastic supports, silicon supports, chips, filters, films, microplates, slides, plastic materials including porous plates or films (e.g., formed by polyethylene, polypropylene, polyamide, polyvinylidene fluoride), and / or wafers, combs, needles or pinheads (e.g., suitable for combined synthesis or analysis of needle arrays) or beads, flat surfaces such as recesses or nanoliter well arrays of wafers (e.g., silicon wafers), wafers with recesses (with or without filter bottoms).

[0310] The solid support may include a polymer matrix (e.g., a gel, a hydrogel). The polymer matrix may be capable of permeating intracellular spaces (e.g., around organelles). The polymer matrix may be capable of being pumped throughout the circulatory system.

[0311] Substrate and microwell array

[0312] As used herein, substrate can refer to a solid support type. Substrate can refer to a solid support that can contain a barcode or a random barcode of the present disclosure. Substrate can, for example, include more than one microwell. Substrate can, for example, be a hole array including two or more microwells. In some embodiments, microwells can include a small reaction chamber of a defined volume. In some embodiments, microwells can capture one or more cells. In some embodiments, microwells can only capture one cell. In some embodiments, microwells can capture one or more solid supports. In some embodiments, microwells can only capture one solid support. In some embodiments, microwells capture single cells and single solid supports (e.g., beads). Microwells can contain barcode reagents of the present disclosure.

[0313] Barcoding method

[0314] The present disclosure provides a method for estimating the number of different targets at different locations in a body sample (e.g., tissue, organ, tumor, cell). The method may include placing a barcode (e.g., a random barcode) very close to the sample, lysing the sample, associating different targets with the barcode, amplifying the target and / or digitally counting the target. The method may also include analyzing and / or visualizing the information obtained from the spatial markers on the barcode. In some embodiments, the method includes visualizing more than one target in the sample. Mapping more than one target to a map of the sample may include generating a two-dimensional map or a three-dimensional map of the sample. Two-dimensional maps and three-dimensional maps may be generated before or after barcoding more than one target in the sample (e.g., random barcoding). Visualizing more than one target in the sample may include mapping more than one target to a map of the sample. Mapping more than one target to a map of the sample may include generating a two-dimensional map or a three-dimensional map of the sample. Two-dimensional maps and three-dimensional maps may be generated before or after barcoding more than one target in the sample. In some embodiments, the two-dimensional map and the three-dimensional map can be generated before or after the sample is lysed. Lysing the sample before or after generating the two-dimensional map or the three-dimensional map can include heating the sample, contacting the sample with a detergent, changing the pH of the sample, or any combination thereof.

[0315] In some embodiments, barcoding more than one target comprises hybridizing more than one barcode to more than one target to generate barcoded targets (e.g., randomly barcoded targets). Barcoding more than one target can comprise generating an indexed library of barcoded targets. Generating an indexed library of barcoded targets can be performed with a solid support comprising more than one barcode (e.g., random barcodes).

[0316] Bring the sample and barcode into contact

[0317] The present disclosure provides methods for contacting a sample (e.g., a cell) with a substrate of the present disclosure. Samples including, for example, thin sections of cells, organs, or tissues can be contacted with bar codes (e.g., random bar codes). Cells can be contacted, for example, by gravity flow, wherein the cells can be precipitated and a monolayer can be produced. The sample can be a thin section of tissue. The thin section can be placed on a substrate. The sample can be one-dimensional (e.g., forming a flat surface). The sample (e.g., a cell) can be dispersed throughout a substrate, for example, by growing / culturing cells on a substrate.

[0318] When the barcode is in close proximity to the target, the target can hybridize to the barcode. The barcodes can be contacted in a non-exhaustive ratio so that each different target can associate with a different barcode of the present disclosure. To ensure effective association between the target and the barcode, the target and the barcode can be cross-linked.

[0319] Cell lysis

[0320] After the distribution of cells and barcodes, the cells can be lysed to release the target molecules. Cell lysis can be accomplished by any of a variety of means, such as by chemical or biochemical means, by osmotic shock, or by means of thermal lysis, mechanical lysis or optical lysis. Cells can be lysed by adding a cell lysis buffer containing a detergent (e.g., SDS, lithium dodecyl sulfate, Triton X-100, Tween-20 or NP-40), an organic solvent (e.g., methanol or acetone) or a digestive enzyme (e.g., proteinase K, pepsin or trypsin) or any combination thereof. In order to increase the association of the target with the barcode, the diffusion rate of the target molecule can be changed by, for example, lowering the temperature of the lysate and / or increasing the viscosity of the lysate.

[0321] In some embodiments, filter paper can be used to crack the sample. The filter paper can be soaked with the lysis buffer on the filter paper top. The filter paper can be applied to the sample with pressure, which can promote the cracking of the sample and the hybridization of the target and substrate of the sample.

[0322] In some embodiments, cleavage can be performed by mechanical cleavage, thermal cleavage, optical cleavage and / or chemical cleavage. Chemical cleavage can include the use of digestive enzymes such as proteinase K, pepsin and trypsin. Cleavage can be performed by adding lysis buffer to the substrate. The lysis buffer can include Tris HCl. The lysis buffer can include at least about 0.01M, 0.05M, 0.1M, 0.5M or 1M or more Tris HCl. The lysis buffer can include up to about 0.01M, 0.05M, 0.1M, 0.5M or 1M or more Tris HCl. The lysis buffer can include about 0.1M Tris HCl. The pH of the lysis buffer can be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or higher. The pH of the lysis buffer can be up to about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or higher. In some embodiments, the pH of the lysis buffer is about 7.5. The lysis buffer can include salt (e.g., LiCl). The salt concentration in the lysis buffer can be at least about 0.1M, 0.5M, or 1M or more. The salt concentration in the lysis buffer can be up to about 0.1M, 0.5M, or 1M or more. In some embodiments, the salt concentration in the lysis buffer is about 0.5M. The lysis buffer can contain a detergent (e.g., SDS, lithium dodecyl sulfate, triton X, tween, NP-40). The detergent concentration in the lysis buffer can be at least about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6% or 7% or more. The detergent concentration in the lysis buffer can be up to about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6% or 7% or more. In some embodiments, the detergent concentration in the lysis buffer is about 1% lithium dodecyl sulfate. The time used in the lysis method can depend on the amount of detergent used. In some embodiments, the more detergent is used, the less time is required for lysis. The lysis buffer can contain a chelating agent (e.g., EDTA, EGTA). The chelating agent concentration in the lysis buffer can be at least about 1mM, 5mM, 10mM, 15mM, 20mM, 25mM or 30mM or more. The chelating agent concentration in the lysis buffer can be at most about 1mM, 5mM, 10mM, 15mM, 20mM, 25mM or 30mM or higher. In some embodiments, the chelating agent concentration in the lysis buffer is about 10mM. The lysis buffer can contain a reducing agent (e.g., β-mercaptoethanol, DTT). The reducing agent concentration in the lysis buffer can be at least about 1mM, 5mM, 10mM, 15mM or 20mM or higher.The reducing agent concentration in the lysis buffer can be up to about 1 mM, 5 mM, 10 mM, 15 mM or 20 mM or higher. In some embodiments, the reducing agent concentration in the lysis buffer is about 5 mM. In some embodiments, the lysis buffer can comprise about 0.1 M TrisHCl, about pH 7.5, about 0.5 M LiCl, about 1% lithium dodecyl sulfate, about 10 mM EDTA and about 5 mM DTT.

[0323] Lysis can be carried out at a temperature of about 4°C, 10°C, 15°C, 20°C, 25°C or 30°C. Lysis can be carried out for about 1 minute, 5 minutes, 10 minutes, 15 minutes or 20 minutes or more minutes. Lysed cells can contain at least about 100,000, 200,000, 300,000, 400,000, 500,000, 600,000 or 700,000 or more target nucleic acid molecules. Lysed cells can contain up to about 100,000, 200,000, 300,000, 400,000, 500,000, 600,000 or 700,000 or more target nucleic acid molecules.

[0324] Attaching barcodes to target nucleic acid molecules

[0325] After cell lysis and release of nucleic acid molecules therefrom, the nucleic acid molecules can be randomly associated with the barcodes of the co-localized solid support. The association can include hybridizing the target recognition region of the barcode with the complementary portion of the target nucleic acid molecule (e.g., the oligo(dT) of the barcode can interact with the poly(A) tail of the target). Assay conditions for hybridization (e.g., buffer pH, ionic strength, temperature, etc.) can be selected to promote the formation of specific stable hybrids. In some embodiments, nucleic acid molecules released from lysed cells can be associated with more than one probe on a substrate (e.g., hybridized with probes on a substrate). When the probe comprises oligo(dT), the mRNA molecule can be hybridized with the probe and reverse transcribed. The oligo(dT) portion of the oligonucleotide can act as a primer for the first strand synthesis of a cDNA molecule. For example, in Fig.12 In a non-limiting example of barcoding shown at block 1216, an mRNA molecule can be hybridized to a barcode on a bead. For example, a single-stranded nucleotide fragment can be hybridized to a target binding region of a barcode.

[0326] Attachment can also include connecting the target recognition region of the barcode to a portion of the target nucleic acid molecule. For example, the target binding region can include a nucleic acid sequence that can be specifically hybridized with a restriction site overhang (e.g., an EcoRI sticky end overhang). The assay procedure can also include treating the target nucleic acid with a restriction enzyme (e.g., EcoRI) to produce a restriction site overhang. The barcode can then be connected to any nucleic acid molecule comprising a sequence complementary to the restriction site overhang. A ligase (e.g., T4 DNA ligase) can be used to connect two fragments. In some embodiments provided herein, the assay procedure can include treating the target nucleic acid with a transposase (e.g., Tn5) to produce a tag fragmentation product. The tag fragmentation product can include an overhang. The tag fragmentation product can be captured by the oligonucleotide barcode provided herein (e.g., by combining the overhang). The tag fragmentation product can include a vacancy. DNA polymerase (e.g., Klenow) can be used to fill the vacancy before the ligation reaction.

[0327] For example, in Fig.12 In the non-limiting example of barcoding shown at block 1220, labeled targets (e.g., target barcode molecules) from more than one cell (or more than one sample) can then be pooled, for example, into a tube. The labeled targets can be pooled by, for example, retrieving barcodes and / or beads to which target barcode molecules are attached.

[0328] Recovery of a solid support-based collection of attached target barcode molecules can be achieved by using magnetic beads and an externally applied magnetic field. After pooling the target barcode molecules, all further processing can be performed in a single reaction vessel. Further processing can include, for example, reverse transcription reactions, amplification reactions, cleavage reactions, dissociation reactions, and / or nucleic acid extension reactions. Further processing reactions can be performed within the microwells, i.e., without first pooling labeled target nucleic acid molecules from more than one cell.

[0329] Reverse transcription or nucleic acid extension

[0330] The present disclosure provides methods for using reverse transcription (e.g., Fig.12The method of producing a target-barcode conjugate by a barcode or nucleic acid extension. The target barcode conjugate may include a barcode and a complementary sequence of all or a portion of the target nucleic acid (i.e., a barcoded cDNA molecule, such as a randomly barcoded cDNA molecule). Reverse transcription of the associated RNA molecule may occur by adding a reverse transcription primer together with a reverse transcriptase. The reverse transcription primer may be an oligo(dT) primer, a random hexanucleotide primer, or a target-specific oligonucleotide primer. The length of the oligo(dT) primer may be 12-18 nucleotides or may be about 12-18 nucleotides and bind to the endogenous poly(A) tail at the 3' end of the mammalian mRNA. The random hexanucleotide primer may bind to the mRNA at each complementary site. The target-specific oligonucleotide primer typically selectively triggers the mRNA of interest.

[0331] In some embodiments, the reverse transcription of mRNA molecules to the RNA molecules of the mark can occur by adding a reverse transcription primer. In some embodiments, the reverse transcription primer is an oligo (dT) primer, a random hexanucleotide primer or a target-specific oligonucleotide primer. Typically, the length of the oligo (dT) primer is 12-18 nucleotides, and is combined with the endogenous poly (A) tail at the 3' end of mammalian mRNA. Random hexanucleotide primers can be combined with mRNA at each complementary site. Target-specific oligonucleotide primers selectively trigger mRNA of interest usually.

[0332] In some embodiments, the target is a cDNA molecule. For example, an mRNA molecule can be reverse transcribed using a reverse transcriptase such as Moloney murine leukemia virus (MMLV) reverse transcriptase to produce a cDNA molecule with a poly (dC) tail. The barcode can include a target binding region with a poly (dG) tail. After base pairing between the poly (dG) tail of the barcode and the poly (dC) tail of the cDNA molecule, the reverse transcriptase converts the template strand from the cellular RNA molecule to the barcode and continues to replicate to the 5' end of the barcode. By doing so, the resulting cDNA molecule contains a barcode sequence (such as a molecular marker) on the 3' end of the cDNA molecule.

[0333] Reverse transcription can occur repeatedly to produce more than one labeled cDNA molecule. The methods disclosed herein may include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 reverse transcription reactions. The methods may include performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 reverse transcription reactions.

[0334] Amplification

[0335] One or more nucleic acid amplification reactions can be performed (e.g., Fig.12 The amplification reaction may be performed in a multiplex manner, wherein more than one target nucleic acid sequence is amplified simultaneously. The amplification reaction may be used to add sequencing adapters to the nucleic acid molecules. The amplification reaction may include amplifying at least a portion of the sample label (if present). The amplification reaction may include amplifying at least a portion of a cell label and / or a barcode sequence (e.g., a molecular marker). The amplification reaction may include amplifying at least a portion of a sample label, a cell label, a spatial label, a barcode sequence (e.g., a molecular marker), a target nucleic acid, or a combination thereof. The amplification reaction can include amplifying 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 100% or a range or number between any two of these values ​​of more than one nucleic acid. The method can also include performing one or more cDNA synthesis reactions to generate one or more cDNA copies of a target barcode molecule comprising a sample marker, a cell marker, a spatial marker, and / or a barcode sequence (e.g., a molecular marker).

[0336] In some embodiments, polymerase chain reaction (PCR) can be used for amplification. As used herein, PCR can refer to a reaction for simultaneously extending the primers of the complementary strands of DNA to amplify a specific DNA sequence in vitro. As used herein, PCR can encompass derivative forms of reactions, including but not limited to RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, and assembly PCR.

[0337] The amplification of the nucleic acid of labeling can include non-PCR-based methods.The example of the non-PCR-based method includes but is not limited to multiple displacement amplification (MDA), transcription-mediated amplification (TMA), amplification based on nucleic acid sequence (NASBA), chain displacement amplification (SDA), real-time SDA, rolling circle amplification or circle to circle amplification (circle-to-circle amplification).Other non-PCR-based amplification methods include more than one cycle of DNA synthesis and transcription of DNA transcription amplification driven by DNA-dependent RNA polymerase or RNA guidance to amplify DNA or RNA targets, ligase chain reaction (LCR) and Qβ replicase (Qβ) method, palindromic probe use, chain displacement amplification, oligonucleotide-driven amplification using restriction endonucleases, primers are hybridized with nucleic acid sequences and the amplification method of the resulting duplex being cracked before extension reaction and amplification, chain displacement amplification, rolling circle amplification and branch extension amplification (RAM) using nucleic acid polymerases lacking 5' exonuclease activity.In some embodiments, amplification does not produce circularized transcripts.

[0338] In some embodiments, the method disclosed herein also includes carrying out polymerase chain reaction to the labeled nucleic acid (e.g., labeled RNA, labeled DNA, labeled cDNA) to produce labeled amplicon (e.g., randomly labeled amplicon). The labeled amplicon can be a double-stranded molecule. The double-stranded molecule can include a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized with a DNA molecule. One or both chains of the double-stranded molecule can include a sample label, a spatial label, a cell label, and / or a barcode sequence (e.g., a molecular label). The labeled amplicon can be a single-stranded molecule. The single-stranded molecule can include DNA, RNA, or a combination thereof. The nucleic acid of the present disclosure can include a synthetic or altered nucleic acid.

[0339] Amplification can include the use of one or more non-natural nucleotides. Non-natural nucleotides can include light unstable or triggerable nucleotides. Examples of non-natural nucleotides can include but are not limited to peptide nucleic acids (PNA), morpholinos and locked nucleic acids (LNA) and glycol nucleic acids (GNA) and threose nucleic acids (TNA). Non-natural nucleotides can be added to one or more cycles of amplified reactions. The interpolation of non-natural nucleotides can be used to identify the product of a specific cycle or time point in an amplified reaction.

[0340] Carrying out one or more amplification reactions can include using one or more primers.One or more primers can include, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 or more nucleotides.One or more primers can include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 or more nucleotides.One or more primers can include less than 12-15 nucleotides.One or more primers can anneal to at least a portion of more than one labeled target (for example, a randomly labeled target).One or more primers can anneal to 3' or 5' ends of more than one labeled target.One or more primers can anneal to the internal region of more than one labeled target. The internal region may be at least about 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 500, 511, 521, 531, 540, 550, 560, 570, 580, 590, 500, 512, 5 In some embodiments, the present invention provides at least one or more primers of the present invention.The present invention provides at least one or more primers of the present invention.The present invention provides at least one or more primers of the present invention.The present invention provides at least one or more primers of the present invention.The present invention provides at least one or more primers of the present invention.The present invention provides at least one or more primers of the present invention.

[0341] One or more primers may include universal primers. Universal primers may anneal to universal primer binding sites. One or more custom primers may anneal to a first sample marker, a second sample marker, a spatial marker, a cell marker, a barcode sequence (e.g., a molecular marker), a target, or any combination thereof. One or more primers may include universal primers and custom primers. Custom primers may be designed to amplify one or more targets. The target may include a subset of total nucleic acids in one or more samples. The target may include a subset of total labeled targets in one or more samples. One or more primers may include at least 96 or more custom primers. One or more primers may include at least 960 or more custom primers. One or more primers may include at least 9600 or more custom primers. One or more custom primers may anneal to two or more different labeled nucleic acids. Two or more different labeled nucleic acids may correspond to one or more genes.

[0342] Any amplification scheme can be used in the methods of the present disclosure. For example, in one approach, the first round of PCR can use gene-specific primers and primers for universal Illumina sequencing primer 1 sequences to amplify molecules attached to beads. The second round of PCR can use nested gene-specific primers flanked by Illumina sequencing primer 2 sequences and primers for universal Illumina sequencing primer 1 sequences to amplify the first PCR product. The third round of PCR adds P5 and P7 and a sample index to turn the PCR product into an Illumina sequencing library. Sequencing using 150bp×2 sequencing can reveal cell markers and barcode sequences (e.g., molecular markers) on read 1, genes on read 2, and sample indexes on index 1 reads.

[0343] In some embodiments, chemical cleavage can be used to remove nucleic acid from substrate. For example, chemical groups or modified bases present in nucleic acid can be used to promote the removal of nucleic acid from solid support. For example, enzymes can be used to remove nucleic acid from substrate. For example, nucleic acid can be removed from substrate by restriction endonuclease digestion. For example, nucleic acid containing dUTP or ddUTP can be removed from substrate by treating nucleic acid with uracil-d-glycosylase (UDG). For example, nucleic acid can be removed from substrate using enzymes (such as, base excision repair enzymes, such as, apurinic / apyrimidinic (ap) endonucleases) that perform nucleotide excision. In some embodiments, photocleavable groups and light can be used to remove nucleic acid from substrate. In some embodiments, cleavable joints can be used to remove nucleic acid from substrate. For example, cleavable joints can include at least one of the following: biotin / avidin, biotin / streptavidin, biotin / neutravidin, Ig protein A, light-labile joints, acid or base-labile joint groups or adapters.

[0344] When the probe is gene specific, the molecule can be hybridized to the probe and reverse transcribed and / or amplified. In some embodiments, after the nucleic acid has been synthesized (e.g., reverse transcribed), it can be amplified. Amplification can be performed in a multiplex manner, wherein more than one target nucleic acid sequence is amplified simultaneously. Amplification can add sequencing adapters to the nucleic acid.

[0345] In some embodiments, amplification can be performed on substrates, for example, with bridging amplification.cDNA can be added with homopolymer tails to produce compatible ends for bridging amplification using oligo (dT) probes on substrates.In bridging amplification, the primer complementary to the 3' end of the template nucleic acid can be the first primer of each pair of primers covalently attached to solid particles.When the sample containing the template nucleic acid contacts the particle and performs a single thermal cycle, the template molecule can be annealed to the first primer, and the first primer is extended forward by adding nucleotides to form a duplex molecule, which is composed of the template molecule and the newly formed DNA chain complementary to the template.In the heating step of the next cycle, the duplex molecule can be denatured, the template molecule is released from the particle, and the complementary DNA chain attached to the particle by the first primer is left.In the annealing stage of the subsequent annealing and extension steps, the complementary chain can be hybridized with the second primer, and the second primer is complementary to the segment of the complementary chain at the position removed from the first primer. This hybridization can cause complementary strands to form bridges between the first primer and the second primer, connected by covalent bonds (secure to) the first primer and connected by hybridization to the second primer. In the extension phase, in the same reaction mixture, by adding nucleotides, the second primer can be extended in the opposite direction, thereby converting the bridge into a double-stranded bridge. Then start the next cycle, and the double-stranded bridge can be denatured to produce two kinds of single-stranded nucleic acid molecules, and one end of each kind of single-stranded nucleic acid molecule is attached to the particle surface via the first primer and the second primer respectively, and the other end of each kind of single-stranded nucleic acid molecule is unattached. In the annealing and extension steps of this second cycle, each chain can be hybridized with other complementary primers previously unused on the same particle, to form a new single-stranded bridge. Two previously unused primers of hybridization are now extended so that two new bridges are converted into double-stranded bridges.

[0346] The amplification reaction can include amplifying at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 100% of more than one nucleic acid.

[0347] The amplification of the labeled nucleic acid may include a PCR-based method or a non-PCR-based method. The amplification of the labeled nucleic acid may include an exponential amplification of the labeled nucleic acid. The amplification of the labeled nucleic acid may include a linear amplification of the labeled nucleic acid. Amplification may be performed by polymerase chain reaction (PCR). PCR may refer to a reaction for simultaneously extending the primers of the complementary strands of DNA to amplify a specific DNA sequence in vitro. PCR may encompass derivative forms of reactions, including but not limited to, RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, inhibition PCR, semi-inhibition PCR, and assembly PCR.

[0348] In some embodiments, the amplification of the labeled nucleic acid includes a non-PCR-based method. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), chain displacement amplification (SDA), real-time SDA, rolling circle amplification, or circle-to-circle amplification. Other non-PCR-based amplification methods include more than one cycle of DNA synthesis and transcription of DNA-dependent RNA polymerase-driven RNA transcription amplification or RNA-guided DNA synthesis and transcription to amplify DNA or RNA targets, ligase chain reaction (LCR), Qβ replicase (Qβ), palindromic probes, chain displacement amplification, oligonucleotide-driven amplification using restriction endonucleases, primers hybridized to nucleic acid sequences and the resulting duplexes cleaved before extension reactions and amplification, chain displacement amplification, rolling circle amplification, and branch extension amplification (RAM) using nucleic acid polymerases lacking 5' exonuclease activity.

[0349] In some embodiments, the method disclosed herein also includes performing a nested polymerase chain reaction on the amplified amplicon (e.g., target). The amplicon can be a double-stranded molecule. The double-stranded molecule can include a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized with a DNA molecule. One or both chains of the double-stranded molecule can include a sample label or a molecular identifier tag. Alternatively, the amplicon can be a single-stranded molecule. The single-stranded molecule can include DNA, RNA, or a combination thereof. The nucleic acid of the present invention can include a synthetic or altered nucleic acid.

[0350] In some embodiments, the method includes repeating the amplification of the labeled nucleic acid to produce more than one amplicon. The method disclosed herein can include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amplification reactions. Alternatively, the method includes performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 amplification reactions.

[0351] Amplification can also include adding one or more control nucleic acids to one or more samples comprising more than one nucleic acid. Amplification can also include adding one or more control nucleic acids to more than one nucleic acid. The control nucleic acid can include a control label.

[0352] Amplification can include the use of one or more non-natural nucleotides. Non-natural nucleotides can include light unstable and / or triggerable nucleotides. Examples of non-natural nucleotides include but are not limited to peptide nucleic acids (PNA), morpholinos and locked nucleic acids (LNA) and glycol nucleic acids (GNA) and threose nucleic acids (TNA). Non-natural nucleotides can be added to one or more cycles of an amplified reaction. The interpolation of non-natural nucleotides can be used to identify the product of a specific cycle or time point in an amplified reaction.

[0353] Carrying out one or more amplification reactions can include using one or more primers. One or more primers can include one or more oligonucleotides. One or more oligonucleotides can contain at least about 7-9 nucleotides. One or more oligonucleotides can contain less than 12-15 nucleotides. One or more primers can anneal to at least a portion of more than one labeled nucleic acid. One or more primers can anneal to the 3' end and / or 5' end of more than one labeled nucleic acid. One or more primers can anneal to the internal region of more than one labeled nucleic acid. The internal region can be at least about 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 500, 511, 521, 531, 542, 550, 560, 570, 580, 590, 500, 512, 5 10, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900 or 1000 nucleotides. One or more primers may include a set of fixed primers. One or more primers may include at least one or more custom primers. One or more primers may include at least one or more control primers. One or more primers may include at least one or more housekeeping gene primers. One or more primers may include universal primers. Universal primers may anneal to universal primer binding sites. One or more custom primers can anneal to the first sample tag, the second sample tag, a molecular identifier label, a nucleic acid or their products. One or more primers can include universal primers and custom primers. Custom primers can be designed to amplify one or more target nucleic acids. Target nucleic acids can include a subset of total nucleic acids in one or more samples. In some embodiments, primers are probes attached to an array of the present disclosure.

[0354] In some embodiments, barcoding more than one target in a sample (e.g., random barcoding) further comprises generating an index library of barcoded targets (e.g., randomly barcoded targets) or barcoded fragments of targets. The barcode sequences of different barcodes (e.g., molecular markers of different random barcodes) can be different from each other. Generating an index library of barcoded targets comprises generating more than one index polynucleotide from more than one target in a sample. For example, for an index library of barcoded targets comprising a first index target and a second index target, the marker region of the first index polynucleotide can differ from the marker region of the second index polynucleotide by less than, about less than, at least less than, or at most less than: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 nucleotides, or a number or range of nucleotides between any two of these values. In some embodiments, generating an index library of barcoded targets includes contacting more than one target (e.g., mRNA molecules) with more than one oligonucleotide comprising a poly (T) region and a tag region; and performing first-strand synthesis using a reverse transcriptase to generate single-stranded labeled cDNA molecules (each comprising a cDNA region and a tag region), wherein the more than one target includes at least two mRNA molecules of different sequences, and the more than one oligonucleotide includes at least two oligonucleotides of different sequences. Generating an index library of barcoded targets may also include amplifying single-stranded labeled cDNA molecules to generate double-stranded labeled cDNA molecules; and performing nested PCR on double-stranded labeled cDNA molecules to generate labeled amplicons. In some embodiments, the method may include generating adapter-tagged amplicons.

[0355] Barcoding (e.g., random barcoding) can include the use of nucleic acid barcodes or tags to label individual nucleic acid (e.g., DNA or RNA) molecules. In some embodiments, it includes adding DNA barcodes or tags to cDNA molecules when generating cDNA molecules from mRNA. Nested PCR can be performed to minimize PCR amplification bias. Adapters for use in sequencing (e.g., next generation sequencing (NGS)) can be added. For example, in Fig.12 At block 1232 of the method, the sequencing results can be used to determine the sequence of cellular markers, molecular markers, and nucleotide fragments of one or more copies of the target.

[0356] Fig.131 is a schematic diagram showing a non-limiting exemplary process for generating an index library of barcoded targets (e.g., random barcoded targets), such as an index library of barcoded mRNAs or fragments thereof. As shown in step 1, the reverse transcription process can encode each mRNA molecule with a unique molecular marker sequence, a cell marker sequence, and a universal PCR site. Specifically, by hybridizing (e.g., randomly hybridizing) a set of barcodes (e.g., random barcodes) 1310 with a poly (A) tail region 1308 of an RNA molecule 1302, the RNA molecule 1302 can be reverse transcribed to produce a labeled cDNA molecule 1304 (including a cDNA region 1306). Each of the barcodes 1310 can include a target binding region, such as a poly (dT) region 1312, a marker region 1314 (e.g., a barcode sequence or molecule), and a universal PCR region 1316.

[0357] In some embodiments, the cell marker sequence may comprise 3 to 20 nucleotides. In some embodiments, the molecular marker sequence may comprise 3 to 20 nucleotides. In some embodiments, each of more than one random barcode further comprises one or more of a universal marker and a cell marker, wherein the universal marker is the same for more than one random barcode on the solid support, and the cell marker is the same for more than one random barcode on the solid support. In some embodiments, the universal marker may comprise 3 to 20 nucleotides. In some embodiments, the cell marker comprises 3 to 20 nucleotides.

[0358] In some embodiments, the labeling region 1314 may include a barcode sequence or molecular label 1318 and a cell label 1320. In some embodiments, the labeling region 1314 may include one or more of a universal label, a dimensional label, and a cell label. The length of the barcode sequence or molecular label 1318 may be below, may be about below, may be at least below, or may be at most below: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides or a number or range of nucleotides between any of these values. The length of the cell label 1320 can be, can be about, can be at least, or can be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides, or a number or range of nucleotides between any of these values. The length of the universal label can be, can be about, can be at least, or can be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides, or a number or range of nucleotides between any of these values. The universal label can be the same for more than one random barcode on the solid support, and the cell label is the same for more than one random barcode on the solid support. The length of a dimensional marker can be less than, can be about less than, can be at least less than, or can be at most less than: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides, or a number or range of nucleotides between any of these values.

[0359] In some embodiments, the marker region 1314 can include, include about, include at least, or include at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 different markers or a number or range of different markers between any of these values, such as barcode sequences or molecular markers 1318 and cellular markers 1320. The length of each marker can be, can be about, can be at least, or can be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides, or a number or range of nucleotides between any of these values. A set of barcodes or random barcodes 1310 can include, can be about, can be at least, or can be at most 10, 20, 40, 50, 70, 80, 90, 100 nucleotides. 2 10 3 10 4 10 5 10 6 10 7 10 8 10 9 10 10 10 11 10 12 10 13 10 14 10 15 10 20 The barcodes or random barcodes 1310 may be a number or range of barcodes or random barcodes 1310 between any of these values. And the group of barcodes or random barcodes 1310 may, for example, each contain a unique tag region 1314. The labeled cDNA molecules 1304 may be purified to remove excess barcodes or random barcodes 1310. Purification may include Ampure bead purification.

[0360] As shown in step 2, the products from the reverse transcription process in step 1 can be pooled into 1 tube and PCR amplified using the first PCR primer pool and the first universal PCR primer. Pooling is possible because of the unique tag region 1314. In particular, the labeled cDNA molecules 1304 can be amplified to produce nested PCR labeled amplicons 1322. Amplification can include multiplex PCR amplification. Amplification can include multiplex PCR amplification with 96 multiplex primers in a single reaction volume. In some embodiments, in a single reaction volume, the multiplex PCR amplification can utilize the following, utilize about the following, utilize at least the following, or utilize at most the following: 10, 20, 40, 50, 70, 80, 90, 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 Amplification may include using a first PCR primer pool 1324 that includes custom primers 1326A-C targeting a specific gene and a universal primer 1328. Custom primers 1326 may hybridize to a region within the cDNA portion 1306' of the labeled cDNA molecule 1304. Universal primers 1328 may hybridize to universal PCR region 1316 of the labeled cDNA molecule 1304.

[0361] like Fig.13As shown in step 3 of , the product from the PCR amplification in step 2 can be amplified with a nested PCR primer pool and a second universal PCR primer. Nested PCR can minimize PCR amplification bias. In particular, the amplicon 1322 of the nested PCR marker can be further amplified by nested PCR. Nested PCR can include multiplex PCR performed in a single reaction volume using a nested PCR primer pool 1330 of nested PCR primers 1332a-c and a second universal PCR primer 1328'. The nested PCR primer pool 1328 can include, include about, include at least, or include at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 different nested PCR primers 1330 or a number or range of different nested PCR primers 1330 between any of these values. Nested PCR primer 1332 may include adapter 1334 and hybridize to a region within cDNA portion 1306″ of labeled amplicon 1322. Universal primer 1328′ may include adapter 1336 and hybridize to universal PCR region 1316 of labeled amplicon 1322. Thus, step 3 produces adapter-tagged amplicon 1338. In some embodiments, nested PCR primer 1332 and second universal PCR primer 1328′ may not include adapter 1334 and adapter 1336. Instead, adapter 1334 and adapter 1336 may be ligated to the product of nested PCR to produce adapter-tagged amplicon 1338.

[0362] As shown in step 4, the PCR products from step 3 can be PCR amplified for sequencing using library amplification primers. In particular, one or more additional assays can be performed on adapter-tagged amplicon 1338 using adapter 1334 and adapter 1336. Adapter 1334 and adapter 1336 can hybridize with primer 1340 and primer 1342. One or more primers 1340 and primer 1342 can be PCR amplification primers. One or more primers 1340 and primer 1342 can be sequencing primers. One or more adapters 1334 and adapter 1336 can be used for further amplification of adapter-tagged amplicon 1338. One or more adapters 1334 and adapter 1336 can be used to sequence adapter-tagged amplicon 1338. Primers 1342 can include a plate index 1344 so that amplicons generated using the same set of barcodes or random barcodes 1310 can be sequenced in one sequencing reaction using next generation sequencing (NGS).

[0363] Binding reagents

[0364] The binding reagents disclosed herein include protein binding reagents. As used herein, "protein binding reagent" has the common and customary meaning understood by those of ordinary skill in the art according to this specification, and refers to a reagent that specifically binds to a protein. For example, a protein binding reagent may include an antibody or a fragment thereof. Examples of suitable protein binding reagents include any protein binding reagent described in U.S. Patent Publication No. 2018 / 0088112 or No. 2018 / 0346970 (each of which is incorporated by reference and incorporated in its entirety with respect to the specific disclosures cited herein). In some embodiments, a binding reagent (e.g., a protein binding reagent) specifically binds to a target protein. In some embodiments, a protein binding reagent includes an antibody that specifically binds to a target protein. As used herein, the term "antibody" has the common and customary meaning understood by those of ordinary skill in the art according to this specification, and refers to a monoclonal antibody, a polyclonal antibody, a multivalent antibody, or a multispecific antibody. Antibodies may include full-length antibodies or binding fragments thereof, such as Fab, Fab', F(ab') 2 , Fab'-SH, Fd, single-chain Fv (scFv), single-chain antibody, disulfide-linked Fv (sdFv), or V L or V H The antibody or its binding fragment can specifically bind to the target protein.

[0365] Exemplary binding agents (e.g., protein binding agents) suitable for the methods, kits, and compositions as described herein can include, for example, antibodies that specifically bind to CTCF (such as antibodies available through Millipore catalog #07-729), H3K27me3 (such as antibodies available through Millipore catalog #07-449, EpiGentek catalog #A4039, or CellSignaling catalog #9733), c-Myc (such as antibodies available through Cell Signaling catalog #D3N8F), Max (such as antibodies available through Santa Cruz catalog #sc-197), Pol II (8WG16; such as antibodies available through Abcam catalog #ab817), H3K4me1 (such as antibodies available through Abcam catalog #ab8895 or EpiGentek catalog #A4031), H3K4me2 (such as antibodies available through Upstate catalog #07-330 or EpiGentek catalog #A4032), H3K4me3 (such as antibodies available through Millipore catalog #05-745, EpiGentek catalog #A4033 or #68393, or Active Motif catalog #39159), H3K27ac (such as antibodies available through Abcam catalog #ab4729, EpiGentek catalog #A4708, or Millipore catalog #MABE647), Oct4 (such as antibodies available through Thermo catalog #701756), Sox2 (such as antibodies available through Active Motif catalog #39843 or Abcam catalog #ab92494), Nanog (such as antibodies available through Active Motif catalog #61419), Brg1 (such as antibodies available through Bethyl Laboratories catalog #A300-813), Suz12 (such as antibodies available through Bethyl Laboratories catalog #A302-407A), IgG (such as mouse IgG available through Millipore catalog #06-371), RNA Pol II Ser2P, RNA Pol II Ser5P, RNA Pol II Ser2+5P, RNA Pol II Ser7P (such as antibodies available through Cell Signaling catalog #54020), NPAT (such as antibodies available through Thermo catalog #PA5-66839), STAT1 (such as antibodies available through Becton Dickinson catalog #610115),NFAT-1 (such as antibodies available through Becton Dickinson catalog #610702), STAT1 (pY701, clone 14 / P-STAT1) (such as antibodies available through Becton Dickinson catalog #612132), STAT1 (pY701, clone 4A) (such as antibodies available through Becton Dickinson catalog #612232), Jun (such as antibodies available through Becton Dickinson catalog #610326), p53 (such as clone PAb240, available through Becton Dickinson catalog #554166), p53 (such as clone 80 / p53, available through Becton Dickinson catalog #610183), Rb (such as antibodies available through Becton Dickinson catalog #554162), STAT3 (such as antibodies available through Becton Dickinson catalog #610189), STAT2 (such as antibodies available through Becton Dickinson catalog #610326), Dickinson catalog #610187), Fos (such as antibodies available through Becton Dickinson catalog #554156), androgen receptor (such as clone G122-434, available through Becton Dickinson catalog #554225), androgen receptor (such as antibody clone G122-25, available through Becton Dickinson catalog #554224), H1K25me1 (such as antibodies available through EpiGentek catalog #A68342), H1K25me2 (such as antibodies available through EpiGentek catalog #A68343), H1K25me3 (such as antibodies available through EpiGentek catalog #A68370), H2(A)K4ac (such as antibodies available through EpiGentek catalog #A68365), H2(A)K5ac (such as antibodies available through EpiGentek catalog #A68350 or A4 300), H2(A)K7ac (such as antibodies available through EpiGentek catalog #A683512), H2(B)K5ac (such as antibodies available through EpiGentek catalog #A68366), H2(B)K12ac (such as antibodies available through EpiGentek catalog #A68352), H2(B)K15ac (such as antibodies available through EpiGentek catalog #A68353), H2(B)K20ac (such as antibodies available through EpiGentek catalog #A68434),H3K4ac (such as antibodies available through EpiGentek catalog #A69361), H3K9ac (such as antibodies available through EpiGentek catalog #A4054, 4022 or A4021), H3K14ac (such as antibodies available through EpiGentek catalog #A4021), H3K18ac (such as antibodies available through EpiGentek catalog #A4024), H3K23ac (such as antibodies available through EpiGentek catalog #A4025), H3K56ac (such as antibodies available through EpiGentek catalog #A68431), H3K6 ... K9me1 (such as antibodies available through EpiGentek catalog #A68435), H3K9me2 (such as antibodies available through EpiGentek catalog #A68435), H3K9me3 (such as antibodies available through EpiGentek catalog #A68435), H3K27me1 (such as antibodies available through EpiGentek catalog #A4037), H3K27me2 (such as antibodies available through EpiGentek catalog #A68392 or A4038), H3K36me (such as antibodies available through EpiGentek catalog #A4040, A68396, A404 1, A4042 or A68388), H3K79me1 (such as antibodies available through EpiGentek catalog #A68419), H3K79me2 (such as antibodies available through EpiGentek catalog #A68419), H3K79me3 (such as antibodies available through EpiGentek catalog #A68419), H4K5ac (such as antibodies available through EpiGentek catalog #A68408, A68356 or A4027), H4K8ac (such as antibodies available through EpiGentek catalog #A4028), H4K12ac (such as antibodies available through #A4029 or A68357), H4K16ac (such as antibodies available through EpiGentek catalog #A68404 or A68398), H4K91ac (such as antibodies available through EpiGentek catalog #A68367), H4K20me (such as antibodies available through EpiGentek catalog #A4046, A4047, A68422 or A4048), or H4K59me (such as antibodies available through EpiGentek catalog #A68337 or A68338), or a combination of two or more of the listed items.

[0366] As used herein, the terms "specific", "specifically" or "specificity" with respect to specific binding of an agent to a target have the ordinary and customary meaning as understood by those of ordinary skill in the art in light of this specification, and refer to a higher or improved binding affinity of an agent to a target compared to the binding affinity of the agent to a non-target. Thus, specific binding refers to preferential binding to the indicated target rather than to a non-target. However, it should be understood that specific binding does not necessarily exclude some minor or insignificant (e.g., background) interactions with substances other than the target. For example, specific binding can refer to a higher or more specific binding affinity than a non-target. -5 The dissociation constant (K) is smaller (indicating tighter binding). D ) combined with, for example, K D Less than 10 -5 M, 10 -6 M, 10 -7 M, 10 -8 M or 10 -9 M. For example, specific binding can mean that the binding affinity of an agent to a target is at least 1.5 times, 2 times, 2.5 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, or 10 times or more greater than the binding affinity of the agent to a non-target. Specific binding can also mean that the binding of an agent to a target is detectably greater, such as at least 1.5 times, 2 times, 2.5 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, or 10 times or more greater than the binding affinity of the agent to a non-target.

[0367] In some embodiments, a binding agent (e.g., a protein binding agent) is associated with a reagent oligonucleotide. In some embodiments, a reagent oligonucleotide includes a unique identifier for a binding agent (e.g., a protein binding agent). As used herein, the term "associated" has the common and customary meanings understood by those of ordinary skill in the art according to this specification, and refers to the interaction between two or more components, and may include covalent or non-covalent bonding. The chemical properties of a covalent bond (two atoms share one or more pairs of valence electrons) are known in the art and include, for example, a disulfide bond or a peptide bond. A non-covalent bond is a chemical bond between atoms or molecules that does not involve a shared valence electron pair. For example, non-covalent interactions include, for example, hydrophobic interactions, hydrogen bond interactions, ionic bonding, van der Waals bonding, or dipole-dipole interactions.

[0368] As used herein, the term "fusion protein" has the common and customary meaning understood by those of ordinary skill in the art according to this specification. It refers to all or part of a polypeptide or compound associated with a DNA digestive enzyme. As described herein, association can be covalent or non-covalent. In some embodiments, the DNA digestive enzyme comprises any enzyme or functional fragment thereof capable of digesting nucleic acid, is essentially composed of any enzyme or functional fragment thereof capable of digesting nucleic acid, or is composed of any enzyme or functional fragment thereof capable of digesting nucleic acid. In some embodiments, the DNA digestive enzyme includes a restriction enzyme, micrococcal nuclease I or Tn5 transposase or its functional fragment. In some embodiments, the fusion protein comprises a domain that specifically binds to a binding agent (e.g., a protein binding agent). In some embodiments, the domain comprises any polypeptide or binding portion thereof that can specifically bind to a binding agent, such as at least one of protein A, protein G, protein A / G or protein L. In some embodiments, the fusion protein is combined with a protein binding agent. In one embodiment, the protein binding agent and the fusion protein are associated as a single complex before use. In another embodiment, the protein binding agent and the fusion protein are not associated as a single complex, but are separate complexes that associate after use (e.g., after the composition enters the nucleus of a cell). In some embodiments, whether the protein binding agent and the fusion protein are associated or separate before use, the size of the protein binding agent, the fusion protein, or the complex including the protein binding agent associated with the fusion protein is sufficient to enter the permeabilized cell and / or the size of the nuclear pore of the cell nucleus.

[0369] Restriction enzymes suitable for the DNA digestion enzymes of the methods, kits and compositions of some embodiments include, for example, type I (e.g., EcoAI, EcoK or EcoB), type II (e.g., HhaI, Hind II, HindIII, BamHI, NotI, NdeI, PacI, PvuI, EcoRI, EcoRII, Sau3AI, SmaI or TaqI), type III (e.g., EcoP15, PstII, PhaBI or HinfIII), type IV (e.g., McrBC or Mrr) and / or type V (e.g., cas9-gRNA complex from CRISPR) restriction enzymes. In some cases, type I enzymes are complex, multi-subunit, combination restriction and modification enzymes that randomly cut DNA away from their recognition sequences. In general, type II enzymes cut DNA at defined positions close to or within their recognition sequences. They can produce discrete restriction fragments and different gel band patterns. Type III enzymes are also large combination restriction and modification enzymes. Type III enzymes typically cleave outside of their recognition sequences and may require two such sequences in opposite orientations within the same DNA molecule to complete cleavage; Type III enzymes rarely give complete digests. In some cases, Type IV enzymes recognize modified (usually methylated) DNA and can be exemplified by the McrBC and Mrr systems of E. coli.

[0370] In some embodiments, the nuclear target is a target protein (e.g., a chromatin-associated protein). As used herein, "target protein" has the common and customary meaning as understood by those of ordinary skill in the art in accordance with this specification, and refers to a polypeptide of interest, and a binding agent (e.g., a protein binding agent) specifically binds to it. In some embodiments, the target protein is associated with the chromatin or DNA of the cell. In some embodiments, the target protein includes any protein known or found to be associated with the chromatin or DNA of the cell. In some embodiments, the target protein is a chromatin-associated protein or a DNA-associated protein. Target proteins associated with chromatin or DNA include, for example, ALC1, androgen receptor, Bmi-1, BRD4, Brg1, coREST, C-jun, c-Myc, CTCF, EED, EZH2, Fos, histone H1, histone H2A, histone H2B, histone H3, histone H4, heterochromatin protein-1γ, heterochromatin protein-1β, HMGN2 / HMG-17, HP1α, HP1γ, hTERT, jun, KLF4, K-Ras, Max, MeCP2, MLL / HRX, NPAT, p300, Nanog, NFAT-1, Oct4, p53, Pol II (8WG16), RNA Pol II Ser2P, RNA Pol II Ser5P, RNA Pol II Ser2+5P, RNA Pol II Ser7P, Rb, RNA polymerase II, SMCI, Sox2, STAT1, STAT2, STAT3, Suz12, Tip60, or UTF1.In some embodiments, the protein binding agent specifically binds to an epitope that includes a methylated (me), phosphorylated (ph), ubiquitinated (ub), ubiquitinated (su), biotinylated (bi), or acetylated (ac) histone residue, including, for example, H1S27ph, H1K25me1, H1K25me2, H1K25me3, H1K26me, H2(A)K4ac, H2(A)K5ac, H2(A)K7ac, H2(A)S1ph, H2 (A)T119ph, H2(A)S122ph, H2(A)S129ph, H2(A)S139ph, H2(A)K119ub, H2(A)K126su, H2(A)K9bi, H2(A)K13 bi, H2(B)K5ac, H2(B)K11ac, H2(B)K12ac, H2(B)K15ac, H2(B)K16ac, H2(B)K20ac, H2(B)S10ph, H2(B)S14ph , H2(B)33ph, H2(B)K120ub, H2(B)K123ub, H3K4ac, H3K9ac, H3K14ac, H3K18ac, H3K23ac, H3K27ac, H3K56ac , H3K4me1, H3K4me2, H3K4me3, H3R8me, H3K9me1, H3K9me2, H3K9me3, H3R17me, H3K27me1, H3K27me2, H3K27me 3, H3K36me, H3K79me1, H3K79me2, H3K79me3, H3K122ac, H3T3ph, H3S10ph, H3T11ph, H3S28ph, H3K4bi, H3K9bi, H3K18bi, H4K5ac, H4K8ac, H4K12ac, H4K16ac, H4K91ac, H4R3me, H4K20me, H4K59me, H4S1ph, H4K12bi or H4 n-terminal tail ubiquitination. Therefore, in some embodiments, the target protein includes any one of the listed proteins, or a combination or two or more of the listed proteins.

[0371] Some embodiments described herein include a first composition comprising a binding agent (eg, a protein binding agent) as described herein and a fusion protein.

[0372] Oligonucleotide probes

[0373] According to the methods, kits and compositions of some embodiments herein, oligonucleotide probes (e.g., oligonucleotide barcodes) are described. In some embodiments, oligonucleotide probes include any oligonucleotide probes described in U.S. Patent Publication No. 2015 / 0299784 or No. 2018 / 0276332 (each of which is incorporated by reference and with its entirety about the specific disclosures cited herein). In the methods, kits and compositions of some embodiments, the second composition comprises a first oligonucleotide probe and a second oligonucleotide probe.

[0374] As used herein, the term "oligonucleotide probe" has the common and customary meaning understood by those of ordinary skill in the art according to this specification, and refers to an oligonucleotide comprising a nucleotide sequence complementary to at least a portion of a target sequence. For example, a nucleotide sequence complementary to at least a portion of a target sequence may include a "capture sequence" or "target binding region" as described herein. As used herein, "complementary" has the common and customary meaning understood by those of ordinary skill in the art according to this specification, and refers to the property that at least a portion of an oligonucleotide can form a hydrogen bond with a target sequence when an oligonucleotide and a target sequence are aligned in reverse. Therefore, complementarity can refer to the ability to accurately pair between two nucleotides. For example, if the nucleotides of a nucleic acid at a given position can form a hydrogen bond with the nucleotides of another nucleic acid, the two nucleic acids are considered to be complementary to each other at that position. The complementarity between two single-stranded nucleic acid molecules can be "partial", in which only some nucleotides are bound, or it can be complete when there is full complementarity between single-stranded molecules. If a first nucleotide sequence is complementary to a second nucleotide sequence, the first nucleotide sequence can be referred to as the "complement" of the second sequence. If a first nucleotide sequence is complementary to a sequence that is opposite to a second sequence (i.e., the nucleotide order is reversed), the first nucleotide sequence may be referred to as the "reverse complement" of the second sequence. As used herein, the terms "complement," "complementary," and "reverse complement" may be used interchangeably. It is understood from this disclosure that if a molecule can hybridize to another molecule, it may be the complement of the molecule to which it hybridizes.

[0375] Complementary nucleobase pairs include adenine (A) and thymine (T), adenine (A) and uracil (U), cytosine (C) and guanine (G), 5-methylcytosine (mC) and guanine (G). Complementary oligonucleotides and / or nucleic acids do not need to have nucleobase complementarity at each nucleoside. On the contrary, some mismatches are tolerated. As used herein, "complete complementarity" or "100% complementarity" related to oligonucleotides means that such oligonucleotides are complementary to another oligonucleotide or nucleic acid at each nucleoside of the oligonucleotide. Complementarity with at least a portion of the target sequence can include complete complementarity, such as 100% complementarity, or less than 100% complementarity, such as 95%, 90%, 80%, 70%, 60%, 50%, 40% or 30% complementarity, as long as complementarity is enough to enable the oligonucleotide probe to be combined with the target sequence.

[0376] In some embodiments, the first oligonucleotide probe comprises a first target binding region (which may also be referred to as a "capture sequence") and a sample identifier sequence. For example, the first target binding region may be placed 3' to the sample identifier sequence. Thus, when the first target binding region is extended along the target hybridization and 5'->3', the strand complementary to the target may be barcoded with the sample identifier sequence. In some embodiments, the first target binding region is complementary to at least a portion of a reagent oligonucleotide associated with a protein binding reagent. In some embodiments, contacting the first oligonucleotide probe with the protein binding reagent results in hybridization of the first target binding region with the reagent oligonucleotide, which may associate the first oligonucleotide probe with the protein binding reagent. It is contemplated that the first target binding region may comprise any suitable sequence that is complementary to the protein binding reagent or a portion thereof and is therefore capable of hybridizing with the protein binding reagent or a portion thereof at the reaction temperature. For example, the first target binding region may comprise at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 consecutive nucleotides that are complementary to the sequence of the reagent oligonucleotide. In some embodiments, the first target binding region (which is also referred to as a "capture sequence") comprises, consists essentially of, or consists of a poly T sequence, and the reagent oligonucleotide comprises a poly A sequence. The poly T sequence may include a sequence of consecutive thymidines of more than 2, 3, 4, 5, 6, 7, 8, 9, or 10 thymidines, and the poly A sequence has a complementary number of consecutive adenines or more than 2, 3, 4, 5, 6, 7, 8, 9, or 10 adenines. It is also contemplated that in some embodiments, the reagent oligonucleotide does not comprise a poly A sequence, and therefore, the first target binding region does not comprise a poly T sequence. In some embodiments, the first target binding region does not include two, three, four, or five consecutive thymidines.

[0377] As used herein, "hybridization" has the ordinary and customary meaning as understood by those of ordinary skill in the art in light of this specification, and refers to the pairing or annealing of complementary oligonucleotides and / or nucleic acids. Although not limited to a particular mechanism, the most common hybridization mechanism involves hydrogen bonding between complementary nucleobases, which may be Watson-Crick, Hoogsteen or reversed Hoogsteen hydrogen bonding.

[0378] In the compositions and / or methods of some embodiments, the second oligonucleotide probe comprises a second target binding region and a sample identifier sequence. In some embodiments, the second target binding region is complementary to at least a portion of a single-stranded overhang of a DNA associated with a target protein as described herein. The single-stranded overhang of the DNA associated with the target protein can be produced by a DNA digestion enzyme, such as a restriction enzyme capable of producing a restriction site overhang, micrococcal nuclease I, or Tn5 transposase.

[0379] In some embodiments, compositions and methods include one or more second oligonucleotide probes. For example, the second oligonucleotide probe can include a plurality of oligonucleotide probes, consist essentially of a plurality of oligonucleotide probes, or consist of a plurality of oligonucleotide probes, particularly for a multiplex comprising more than one antibody. In such embodiments, the Tn5 transposase can include different adapter loading antibodies such that the DNA bound by each antibody is bound to a different adapter. Different oligonucleotide probes complementary to each adapter can be attached to a substrate so that a substrate includes more than two different types of oligonucleotide probes to capture DNA fragments added by different adapters.

[0380] As used herein, the term "sample identifier sequence" has the common and customary meaning understood by those of ordinary skill in the art in light of this specification and refers to a sequence of a first oligonucleotide probe or a second oligonucleotide probe that is used to identify an oligonucleotide probe. The sample identifier sequence can be any sequence that can be used to specifically identify an oligonucleotide probe. In some embodiments, the first oligonucleotide probe and the second oligonucleotide probe include the same sample identifier sequence. In some embodiments, the first oligonucleotide probe includes a first sample identifier sequence, and the second oligonucleotide probe includes a second sample identifier sequence. In some embodiments, the first sample identifier sequence and the second sample identifier sequence are each from a set of different unique sample identifier sequences. The length of the sample identifier sequence can depend on the number (or estimated number) of samples. For example, up to 10 may be present in more than one solid support (e.g., beads). 6In some embodiments, the sample identifier sequence can be a cell marker. ...

[0381] The barcode sequence comprises or consists of a nucleic acid sequence that can be used to identify a nucleic acid, such as target protein-associated DNA from a cell, or an amplicon or reverse transcript derived from target protein-associated DNA from a cell. For example, the barcode sequence can comprise at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, or 50 nucleotides, including ranges between any two of the recited values, such as 4-10, 4-15, 4-20, 4-30, 4-50, 6-10, 6-15, 6-20, 6-30, 6-50, 8-10, 8-15, 8-20, 8-30, or 8-50. In some embodiments, the barcode sequence comprises a random sequence. In some embodiments, there can be at least or at most 10 2 10 3 10 4 10 5 10 6 10 7 10 8 10 9 A unique barcode sequence.

[0382] In some embodiments, the first oligonucleotide probe and the second oligonucleotide probe are fixed on substrate. In some embodiments, substrate comprises any substrate on which the oligonucleotide sequence can be fixed, such as a bead, a microarray, a plate, a pipe or a hole. In some embodiments, the first oligonucleotide probe and the second oligonucleotide probe are fixed to the same substrate. Substrate can be basically composed of or composed of: a bead, a microarray, a plate, a pipe or a hole.

[0383] As used herein, the terms "tethered," "attached," and "immobilized" may be used interchangeably and may refer to covalent or non-covalent means for attaching a barcode to a solid support. Any of a variety of different solid supports may be used as a solid support for immobilizing oligonucleotide probes.

[0384] Substrate can include a type of solid support. Substrate can refer to a continuous solid or semi-solid surface on which the method of the present disclosure can be performed. For example, substrate can refer to an array, a cartridge, a chip, a device, and a slide. Therefore, "solid support" and "substrate" can be used interchangeably.

[0385] Substrate can include or be substantially composed of beads, films, paper, plastics, coated surfaces, flat surfaces, glass, slides, chips or any combination thereof. In some embodiments, at least one surface of the support can be substantially flat, although in some embodiments, it may be desirable to physically separate the synthesis areas for different compounds with, for example, holes, raised areas, needles, etched grooves, etc. According to some embodiments, the substrate can include, be substantially composed of, or be composed of: resins, gels, microspheres or other geometric configurations. According to some embodiments, the substrate includes silica chips, micron particles, nanoparticles, plates and arrays. Solid supports can include beads (e.g., silica gel, controlled aperture glass, magnetic beads, Dynabead, Wang resin; Merrifield resin, Sephadex / agarose gel beads, cellulose beads, polystyrene beads, etc.), capillaries, flat supports such as glass fiber filters, glass surfaces, metal surfaces (steel, gold, silver, aluminum, silicon and copper), glass supports, plastic supports, silicon supports, chips, filters, films, microplates, slides, etc. Plastic materials include beads in an array of recessed or nanoliter wells in a porous plate or membrane (e.g., formed from polyethylene, polypropylene, polyamide, polyvinylidene fluoride), wafers, combs, needles or pins (e.g., needle arrays suitable for combinatorial synthesis or analysis), or a flat surface such as a wafer (e.g., a silicon wafer), a wafer with recesses (with or without a filter bottom).

[0386] The substrate or solid support according to some embodiments herein may include any type of solid, porous or hollow sphere, ball, socket, cylinder or other similar configuration, including plastic, ceramic, metal or polymeric material (e.g., hydrogel), on which nucleic acid can be fixed (e.g., covalently or non-covalently). The substrate or solid support may include discrete particles that may be spherical (e.g., microspheres) or have a non-spherical or irregular shape, such as a cubic, rectangular, conical, cylindrical, conical, elliptical or disc-shaped, etc. More than one solid support spaced apart in an array may not include a substrate. The solid support may be used interchangeably with the term "bead".

[0387] Substrates such as beads may include a variety of materials, including but not limited to paramagnetic materials (e.g., magnesium, molybdenum, lithium, and tantalum), superparamagnetic materials (e.g., ferrites (Fe 3 O 4 ; magnetite) nanoparticles), ferromagnetic materials (e.g., iron, nickel, cobalt, some alloys thereof, and some rare earth metal compounds), ceramics, plastics, glass, polystyrene, silica, methylstyrene, acrylic polymers, titanium, latex, agarose gel, agarose, hydrogels, polymers, cellulose, nylon, or any combination thereof.

[0388] Methods for labeling target protein-associated DNA

[0389] Some embodiments described herein relate to methods for labeling target protein-associated DNA using binding agents (e.g., protein binding agents), fusion proteins, and oligonucleotide probes as described herein. For example, the method may include permeabilizing a cell comprising a target protein, contacting the permeabilized cell with a digestion composition comprising, for example, a fusion protein and a protein binding agent (the protein binding agent is associated with a reagent oligonucleotide) such that the protein binding agent and the fusion protein form a complex. The method may include digesting the target protein-associated DNA in the cell with a DNA digestion enzyme of the fusion protein, thereby forming a single-stranded overhang of the DNA. The method may include contacting the cell with a first oligonucleotide probe and a second oligonucleotide probe, the first oligonucleotide probe and the second oligonucleotide probe hybridizing with the single-stranded overhang of the reagent oligonucleotide and the DNA, respectively. The method may include extending the first oligonucleotide probe and the second oligonucleotide probe. The binding and digestion of the binding agent (e.g., protein binding agent) to the target protein may be performed in situ (e.g., in the nucleus of the cell). In some embodiments, the method may include labeling, sequencing, amplifying, or generating a library of target protein-associated DNA.

[0390] In some embodiments, the method includes permeabilizing the cell. The cell may include any cell including a nucleic acid of interest (such as a DNA associated with a target protein). As used herein, "permeabilization" has the common and customary meanings understood by those of ordinary skill in the art according to this specification, and refers to a process that promotes access to the cell cytoplasm or intracellular molecules, components of the cell or structure. The permeabilization of the cell may include chemical or physical permeabilization, such as exposing the cell to sonication, formaldehyde, ethanol, detergents, surfactants, guanidine hydrochloride, digitonin or Triton X. In some embodiments, the permeabilized cell includes a complete nucleus. In some embodiments, the permeabilized cell includes chromatin that is associated with genomic DNA. In some embodiments, the nucleus is separated, and the kits and methods described herein include contacting the reagents described herein with the separated nucleus. Therefore, it should be understood that anywhere in the description of cells containing nuclei, separated nuclei are also clearly considered. Therefore, anywhere in the description of methods, kits, compositions or uses comprising cells containing nuclei, separated nuclei can be used to replace cells containing nuclei.

[0391] In some embodiments, permeabilized cells are contacted with a composition as described herein, which embodies a binding agent (e.g., a protein binding agent) and a fusion protein (comprising, for example, a portion of a "first composition" as described herein). In some embodiments, the protein binding agent and the fusion protein are sized to enter the permeabilized cell (e.g., diffuse through a permeabilization opening having a diameter greater than the largest diameter of the fusion protein) and are also sized to enter the nucleus of the cell (e.g., diffuse through a nuclear pore).

[0392] In some embodiments, the size of the binding agent (e.g., protein binding agent) and the fusion protein is set to enter the permeabilized cell and the nucleus of the cell (or the separated nucleus) separately, so that the size of the protein binding agent without the fusion protein is set to enter the permeabilized cell and the nucleus of the cell, and the fusion protein without the protein binding agent has a sufficient size to enter the permeabilized cell and the nucleus of the cell. In some embodiments, the size of the protein binding agent and the fusion protein is set to diffuse through the nuclear pores of the cell. In some embodiments, after the permeabilized cell is contacted with the composition, the protein binding agent and the fusion protein enter the permeabilized cell and the nucleus of the cell respectively, and associate with each other after entering the nucleus of the cell. Therefore, in some embodiments, the permeabilized cell is first contacted with the protein binding agent (e.g., by contacting the permeabilized cell with a composition comprising a protein binding agent, substantially consisting of a protein binding agent, or consisting of a protein binding agent), and then the permeabilized cell is contacted with the fusion protein (e.g., by contacting the permeabilized cell with a composition comprising a fusion protein). In some embodiments, before the permeabilized cells are contacted with the composition comprising the protein binding reagent, the composition comprising, consisting essentially of, or consisting of the fusion protein is contacted with the permeabilized cells. In any embodiments described herein, the protein binding reagent and the fusion protein can be contacted with the permeabilized cells separately (e.g., sequentially) so that the protein binding reagent and the fusion protein enter the permeabilized cells and the nucleus of the cells, respectively.

[0393] In some embodiments, binding reagent (for example, protein binding reagent) and fusion protein are mutually associated before contacting with permeabilized cells. In such embodiments, the size of the associated protein binding reagent and fusion protein is set to mutually associate into the permeabilized cell and the nucleus of the cell (for example, through the opening in the cell membrane of the permeabilized cell and / or diffuse through the nuclear pore). In some embodiments, the size of the protein binding reagent associated with the fusion protein is configured to diffuse through the nuclear pore of the cell as a whole. In some embodiments, the combination of the protein binding reagent associated with the fusion protein has the following diameter: no more than 1nm, no more than 2nm, no more than 3nm, no more than 4nm, no more than 5nm, no more than 10nm, no more than 20nm, no more than 30nm, no more than 40nm, no more than 50nm, no more than 60nm, no more than 70nm, no more than 80nm, no more than 90nm, no more than 100nm, no more than 110nm or no more than 120nm, or the size in the range defined by any two values ​​mentioned above. In some embodiments, a composition comprising a protein binding agent is contacted with a composition comprising a fusion protein, and the protein binding agent and the fusion protein are associated with each other, such as by the domain of the specific binding protein binding agent of the fusion protein. In some embodiments, the protein binding agent and the fusion protein are associated with each other. In some embodiments, the associated protein binding agent and the fusion protein are contacted with permeabilized cells as a binding complex, which enters the permeabilized cells and diffuses through the nuclear pores of the cells to bind to the target protein. Therefore, in any embodiment described herein, the protein binding agent combined with the fusion protein can be contacted with the permeabilized cell so that the protein binding agent and the fusion protein are combined and enter the permeabilized cell.

[0394] In some embodiments, a binding agent (e.g., a protein binding agent) binds to a target protein of a permeabilized cell, wherein the target protein is associated with the DNA of the cell. The binding of a protein binding agent to a target protein can include placing a fusion protein adjacent to the DNA associated with the target protein, so that the DNA digestion enzyme fusion protein can interact with the DNA associated with the target protein, and catalyze the digestion of the DNA associated with the target protein. The DNA digestion enzyme of the fusion protein can therefore be brought to the close proximity of the DNA for enzymatic digestion of the DNA. In some embodiments, the DNA digestion enzyme of the fusion protein digests the DNA associated with the target protein. In some embodiments, the digestion of the DNA results in a single-stranded overhang in the digested DNA. In some embodiments, the digestion of the DNA associated with the target protein produces a complex, which includes a fusion protein associated with a protein binding agent bound to the target protein and associated with the digested DNA. The digestion reaction can be quenched or inactivated by subjecting the reaction to a chelating agent (including, for example, EDTA and / or EGTA). In some embodiments, the digestion reaction occurs in two steps, including cleavage and adapter integration.

[0395] In some embodiments, after antibody binding, single cells and beads with oligonucleotides are placed in wells (e.g., microliter wells or nanoliter wells) so that digested DNA from single cells becomes a unique marker. In some embodiments, digestion is triggered by Mg++, Ca++ or enzyme buffer after the single cells are in wells with beads to prevent digested DNA / fusion protein complexes from mixing with those from other cells.

[0396] In some embodiments, as described herein, the formed complex comprising fusion protein, binding reagent (e.g., protein binding reagent), target protein and digested DNA is contacted with a composition comprising the first oligonucleotide probe and the second oligonucleotide probe. Optionally, the complex can be separated from the permeabilized cell before contacting the first oligonucleotide probe and the second oligonucleotide probe. After contact, the first target binding region of the first oligonucleotide probe can be hybridized with the reagent oligonucleotide of the protein binding reagent. The second target binding region of the second oligonucleotide probe can be hybridized with the single-stranded overhang of the digested DNA. The hybridization of the first oligonucleotide probe and the second oligonucleotide probe with their respective complementary nucleic acid sequences therefore causes these respective complementary sequences to associate with the substrate. In some embodiments, the polypeptide portion of the complex is removed (e.g., by contacting with proteinase K), and the nucleic acid is retained. In some embodiments, unbound nucleic acids are removed (e.g., by contacting with nucleases such as ExoI).

[0397] Some embodiments also include further manipulation of the digested DNA, including extension of oligonucleotide probes to produce more than one labeled nucleic acid, capture of complexes on a substrate, ligation, transcription, nucleic acid library construction, sequencing, immunoprecipitation or other analysis. The methods, compositions and kits disclosed herein may also include immunoprecipitation of target protein-associated DNA. Any or all of the methods described herein may be performed using an automated system (such as a thermal cycler or thermal mixer), wherein reverse transcription, ligation, digestion, incubation and / or washing and / or other steps are automated.

[0398] The methods disclosed herein may also include reverse transcription, for example if the RNA or single-stranded DNA is associated with the target protein. Reverse transcription can produce DNA complementary to the RNA or single-stranded DNA. In some cases, at least a portion of the oligonucleotide probe comprises a primer for the reverse transcription reaction.

[0399] One or more nucleic acid amplification reactions can be performed to produce multiple copies of DNA associated with a target protein (target nucleic acid sequence or labeled nucleic acid), for example, to produce a nucleic acid library. Amplification can be performed in a multiplexed manner, wherein more than one target nucleic acid sequence is amplified simultaneously. Amplification can be used to add sequencing adapters to nucleic acid molecules. The amplification reaction can include at least a portion of amplifying sample labels (if present). The amplification reaction can include at least a portion of amplifying cells and / or molecular markers. The amplification reaction can include amplifying at least a portion of the following: sample labels, cell markers, spatial markers, molecular markers, target nucleic acids, or a combination thereof. The amplification reaction can include amplifying at least 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 100% of more than one nucleic acid, or a range or amount between any two of these values. The method can also include performing one or more cDNA synthesis reactions to generate one or more cDNA copies of a target-barcode molecule comprising a sample marker, a cell marker, a spatial marker, and / or a molecular marker.

[0400] As described herein, oligonucleotide probes can include universal labels (which may also be referred to herein as universal primer sites), cell labels, molecular labels, and sample labels, or any combination thereof. For example, oligonucleotide probes can include molecular labels and sample labels. In combination, sample labels can distinguish target nucleic acids between samples, cell labels can distinguish target nucleic acids from different cells in a sample, molecular labels can distinguish different target nucleic acids in a cell (e.g., different copies of the same target nucleic acid), and universal labels can be used to amplify target nucleic acids and sequence target nucleic acids.

[0401] Any of universal markers, molecular markers, cell markers, joint markers and / or sample markers as described herein can comprise a random sequence of nucleotides. The random sequence of nucleotides can be computer generated. The random sequence of nucleotides may not have a pattern associated therewith. Universal markers, molecular markers, cell markers, joint markers and / or sample markers can comprise a non-random (e.g., nucleotides comprising a pattern) sequence of nucleotides. The sequence of universal markers, molecular markers, cell markers, joint markers and / or sample markers can be a commercially available sequence. The sequence of universal markers, molecular markers, cell markers, joint markers and / or sample markers can comprise a random body sequence. A random body sequence can refer to an oligonucleotide sequence of all possible sequences of a random body including a given length. Alternatively or additionally, universal markers, molecular markers, cell markers, joint markers and / or sample markers can comprise a predetermined nucleotide sequence.

[0402] In some embodiments, polymerase chain reaction (PCR) can be used for amplification. As used herein, PCR has its common meaning and can refer to the reaction of a specific DNA sequence for in vitro amplification by primer extension while complementary strands of DNA. As used herein, PCR can include a derivative form of the reaction, including but not limited to RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplexed PCR, digital PCR and assembly PCR.

[0403] The amplification of the DNA associated with the target protein may include a non-PCR-based method. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), amplification based on nucleic acid sequences (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or ring-to-ring amplification. Other non-PCR-based amplification methods include more than one cycle of DNA synthesis and transcription of DNA-dependent RNA polymerase-driven transcription amplification or RNA-guided DNA synthesis and transcription to amplify DNA or RNA targets, ligase chain reaction (LCR) and Qβ replicase (Qβ) methods, the use of palindromic probes, strand displacement amplification, oligonucleotide-driven amplification using restriction endonucleases, primers hybridized to nucleic acid sequences and the resulting duplexes cleaved before extension reactions and amplification, strand displacement amplification, rolling circle amplification, and branch extension amplification (RAM) using nucleic acid polymerases lacking 5' exonuclease activity. In some embodiments, amplification does not produce circularized transcripts.

[0404] Amplification can include the use of one or more non-natural nucleotides. Non-natural nucleotides can include light unstable or triggerable nucleotides. Examples of non-natural nucleotides can include, but are not limited to peptide nucleic acids (PNA), morpholinos and locked nucleic acids (LNA) and glycol nucleic acids (GNA) and threose nucleic acids (TNA). Non-natural nucleotides can be added to one or more cycles of the amplified reaction. Adding non-natural nucleotides can be used to identify the product of a specific cycle or time point in the amplified reaction.

[0405] Amplification as described herein can include the use of one or more primers.One or more primers can each contain, for example, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25 or more nucleotides, including the range between any two values ​​in the listed values.Primers can each contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25 or more nucleotides, including the range between any two values ​​in the listed values.Primers can contain less than 12-15 nucleotides.Primers can each anneal to at least a portion of more than one randomly labeled target.Primers can each anneal to 3' ends or 5' ends (for example, universal primer binding sites described herein) of more than one labeled target. The primers can each anneal to a region, such as an internal region, of more than one labeled target. A region can include at least about 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800 In some embodiments, the present invention provides the primer of the present invention.Primer can comprise one or more custom primers.For example, the 5 ' part of the oligonucleotide probe hybridized with the target can serve as the primer for amplifying the target.Optionally, the primer can comprise at least one or more control primers.Primer can comprise at least one or more gene specific primers.

[0406] One or more primers may include universal primers and / or custom primers. Universal primers may anneal to universal primer binding sites, such as PCR handles as described herein. According to some embodiments, one or more custom primers may anneal to a first sample marker, a second sample marker, a spatial marker, a cell marker, a molecular marker, a target, or any combination thereof. One or more primers may include universal primers and custom primers. Custom primers may be designed to specifically amplify one or more targets. The target may include a subset of total nucleic acids in one or more samples. The target may include a subset of total labeled targets in one or more samples. One or more primers may include at least 96 or more custom primers. One or more primers may include at least 960 or more custom primers. One or more primers may include at least 9600 or more custom primers. One or more custom primers may anneal to two or more different labeled nucleic acids. Two or more different labeled nucleic acids may correspond to one or more genes and / or intergenic nucleic acids.

[0407] Any amplification scheme can be used in the methods of the present disclosure. For example, in one scheme, the first round of PCR can amplify the molecules immobilized on the beads using gene-specific primers and primers for the universal Illumina sequencing primer 1 sequence. The second round of PCR can amplify the first PCR product using nested gene-specific primers flanked by the Illumina sequencing primer 2 sequence and primers for the universal Illumina sequencing primer 1 sequence. The third round of PCR can add P5 and P7 and the sample index to turn the PCR product into an Illumina sequencing library. Sequencing using 150bp×2 sequencing can reveal cell markers and molecular markers on read 1, genes on read 2, and sample indexes on index 1 reads.

[0408] In some embodiments, nucleic acid can be removed from substrate by chemical cleavage. For example, chemical groups or modified bases present in nucleic acid can promote its removal from solid support. For example, enzymes can be used to remove nucleic acid from substrate. For example, nucleic acid can be removed from substrate by restriction endonuclease digestion. For example, nucleic acid containing dUTP or ddUTP can be treated with uracil-d-glycosylase (UDG) to remove nucleic acid from substrate. For example, nucleic acid can be removed from substrate using enzymes (such as, base excision repair enzymes, such as apurinic / apyrimidinic (ap) endonucleases) that perform nucleotide excision. In some embodiments, photocleavable groups and light can be used to remove nucleic acid from substrate. In some embodiments, cleavable joints can be used to remove nucleic acid from substrate. For example, cleavable joints can include at least one of the following: biotin / avidin, biotin / streptavidin, biotin / neutravidin, Ig protein A, light-labile joints, acid or base-labile joint groups or adapters.

[0409] When the oligonucleotide probe is gene-specific, the target molecule can hybridize with the probe and can be reverse transcribed and / or amplified. In some embodiments, after the nucleic acid has been synthesized (e.g., reverse transcribed), the nucleic acid can be amplified. Amplification can be performed in a multiplex manner, wherein multiple target nucleic acid sequences are amplified simultaneously. Amplification can add sequencing adapters to nucleic acids.

[0410] In some embodiments, amplification can be performed on a substrate, for example, by bridging amplification.cDNA can be added with homopolymer tails to produce compatible ends for bridging amplification using oligo (dT) probes on substrates.In bridging amplification, the primer complementary to the 3' end of the template nucleic acid can be the first primer in each pair of primers covalently attached to solid particles.When a sample containing template nucleic acid contacts particles and performs a single thermal cycle, the template molecule can anneal to the first primer and the first primer is extended in the forward direction by adding nucleotides to form a duplex molecule, which comprises a template molecule and a newly formed DNA chain complementary to the template, or is composed of a template molecule and a newly formed DNA chain complementary to the template.In the heating step of the next cycle, the duplex molecule can be denatured, releasing the template molecule from the particle and leaving a complementary DNA chain attached to the particle by the first primer.In the annealing stage of the subsequent annealing and extension step, the complementary chain can be hybridized with the second primer, and the second primer is complementary to the segment of the complementary chain at the position removed from the first primer. This hybridization can cause complementary strands to form a bridge between the first primer and the second primer, connecting the first primer by a covalent bond and connecting the second primer by hybridization. In the extension stage, in the same reaction mixture, by adding nucleotides, the second primer can be extended in the reverse direction, thereby converting the bridge into a double-stranded bridge. Then start the next cycle, and the double-stranded bridge can be denatured to produce two single-stranded nucleic acid molecules, and an end that each single-stranded nucleic acid molecule has is attached to the particle surface via the first primer and the second primer respectively, and the other end of each single-stranded nucleic acid molecule is unattached. In the annealing and extension steps of this second cycle, each chain can be hybridized with other complementary primers that were not previously used on the same particle, to form a new single-stranded bridge. Two previously unused primers that are now hybridized extend thereby converting two new bridges into double-stranded bridges.

[0411] In some embodiments, the method includes repeatedly amplifying the labeled nucleic acid to produce more than one amplicon. The method disclosed herein can include repeating the amplification to perform at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 amplification reactions. Alternatively, the method includes performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or 100 amplification reactions.

[0412] Amplification may also include adding one or more control nucleic acids to one or more samples comprising more than one nucleic acid. Amplification may also include adding one or more control nucleic acids to more than one nucleic acid. The control nucleic acid may include a control marker.

[0413] In some embodiments, the method further comprises sequencing.Sequencing a target nucleic acid (including a labeled target nucleic acid) can include performing a sequencing reaction to determine the sequence of at least a portion of: the target nucleic acid, its complement, its reverse complement, or any combination thereof.

[0414] Determination of the sequence of a labeled target (e.g., an amplified nucleic acid, a labeled nucleic acid, a cDNA copy of a labeled nucleic acid, etc.) can be performed using a variety of sequencing methods, including, but not limited to, sequencing by hybridization (SBH), sequencing by ligation (SBL), quantitative incremental fluorescent nucleotide addition sequencing (QIFNAS), stepwise ligation and cleavage, fluorescence resonance energy transfer (FRET), molecular beacons, TaqMan reporter probe digestion, pyrosequencing, fluorescence in situ sequencing (FISSEQ), FISSEQ beads, wobble sequencing, multiplex sequencing, polymerized colony (POLONY) sequencing; nanogrid rolling circle sequencing (ROLONY), allele-specific oligonucleotide ligation assay (e.g., oligonucleotide ligation assay (OLA), single-template molecule OLA using ligated linear probes and rolling circle amplification (RCA) readout, ligated padlock probes, or single-template molecule OLA using ligated circular padlock probes and rolling circle amplification (RCA) readout), etc.

[0415] In some embodiments, determining the sequence of a target nucleic acid or any product thereof comprises paired end sequencing, nanopore sequencing, high throughput sequencing, shotgun sequencing, dye-terminator sequencing, multi-primer DNA sequencing, primer walking, Sanger dideoxy sequencing, Maxim-Gilbert sequencing, pyrophosphate sequencing, true single molecule sequencing, or any combination thereof. Alternatively, the sequence of a randomly barcoded target or any product thereof can be determined by electron microscopy or a chemically sensitive field effect transistor (chemFET) array.

[0416] High throughput sequencing methods can be utilized, such as cycle array sequencing using platforms such as Roche 454, Illumina Solexa, ABI-SOLiD, ION Torrent, Complete Genomics, Pacific Bioscience, Helicos, or Polonator platforms. In some embodiments, sequencing can include MiSeq sequencing. In some embodiments, sequencing can include HiSeq sequencing.

[0417] Sequencing can include at least about 200, 300, 400, 500, 600, 700, 800, 900, 1,000 or more sequencing reads per run. In some embodiments, sequencing includes sequencing at least or at least about 1500, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000 or 10000 or more sequencing reads per run. Sequencing can include less than or equal to about 1,600,000,000 sequencing reads per run. Sequencing can include less than or equal to about 200,000,000 reads per run.

[0418] Some embodiments of the description provided herein can be understood with reference to the accompanying drawings. Figure 1 , a binding reagent (e.g., protein binding reagent 105) associated with a fusion protein 125 is provided, the fusion protein 125 comprising a digestive enzyme 120 and a domain 130 that specifically binds to the protein binding reagent 105. The protein binding reagent 105 comprises a reagent oligonucleotide 110. In some embodiments, the reagent oligonucleotide 110 comprises a capture sequence, such as a poly A sequence 115. As discussed herein, it is noted that in some embodiments, the capture sequence can be a sequence other than a poly A sequence. In some embodiments, such as Figure 1 As shown, protein binding agent 105 is associated with fusion protein 125 to form associated protein 101. For example, protein binding agent 105 can be non-covalently or covalently bound to fusion protein 125 to form associated protein 101.

[0419] like Figure 2 As shown, the permeabilized cell 205 contains a nuclear target (e.g., a target protein 225) associated with a DNA 215 within the nucleus 210 of the permeabilized cell 205. The nuclear target may include DNA (e.g., methylated DNA) or a nuclear protein (e.g., a gDNA-associated protein). In some embodiments, the DNA 215 may also be associated with other proteins such as an off-target protein 220. In some embodiments, the permeabilized cell 205 is contacted with the associated protein 101, the associated protein 101 enters the permeabilized cell 205 and binds to the target protein 225, forming a permeabilized cell 205 containing the associated protein 101 bound to the target protein 225. In some embodiments, the permeabilized cell 205 is contacted with a protein binding agent 105 and a fusion protein 125, the fusion protein 125 enters the permeabilized cell 205 and forms the associated protein 101. The associated protein can bind to the target protein 225 in the nucleus 210 of the permeabilized cell 205 , forming a permeabilized cell 205 comprising the associated protein 101 bound to the target protein 225 .

[0420] like Figure 3 As shown, the DNA digesting enzyme 120 of the fusion protein 125 digests the DNA 215, and the associated protein 101 forms a complex 305, which includes the protein binding agent 105, the fusion protein 125, the target protein 225, and the digested DNA 215 with a single-stranded overhang. Figure 3 Shown are complexes 305 removed from permeabilized cells 205. In some embodiments, the method comprises cell lysis (eg, by chemical or biochemical means, by osmotic shock, or by thermal, mechanical, or optical lysis) to release the complexes.

[0421] Figure 4 Depicted are a first oligonucleotide probe 405 and a second oligonucleotide probe 410, each affixed to a bead 415. In some embodiments, the first oligonucleotide probe 405 includes a first target binding region 406. In some embodiments, the first target binding region 406 comprises, consists essentially of, or consists of a poly-T sequence. In some embodiments, the second oligonucleotide probe 410 includes a second target binding region 412. In some embodiments, the second oligonucleotide probe 410 comprises, consists essentially of, or consists of a second poly-T sequence 411.

[0422] like Figure 5A As shown, contacting the complex 305 with a substrate 415 including the first oligonucleotide probe 405 and the second oligonucleotide probe 410 results in hybridization of the first target binding region 406 with the reagent oligonucleotide 110, and hybridization of the second target binding region 410 with the single-stranded overhang of the digested DNA 215. Thus, Figure 5A The capture of complex 305 to bead 415 is shown. In such embodiments, material not captured to the bead, including cells or components thereof, is removed, for example, by washing.

[0423] like Figure 5B As shown, in some embodiments, the captured complex can undergo ligation and reverse transcription. Such steps can be performed in an automated instrument (such as a thermal mixer). A ligase (e.g., T4 DNA ligase) can be used to connect the two fragments. The tag fragmentation products produced according to some embodiments provided herein may include gaps. The gaps between the DNA fragments can be filled by a gap filling enzyme (such as a DNA polymerase (e.g., Klenow)). Figure 5C As shown, in some embodiments, proteins, including protein binding agents, fusion proteins, and target proteins are removed. Optionally, polypeptides can be removed, for example, by treating the captured complex with proteinase K. In addition, unbound nucleic acids can be removed, for example, by contact with a nuclease (such as ExoI). Figure 5DAs shown, in some embodiments, the hybridized nucleic acids can undergo further processing to barcode the nucleic acids of the captured complexes (eg, digested DNA and reagent oligonucleotides), including, for example, denaturation, random primer extension, and / or polymerase chain reaction. Figure 6 Depicted is the generation of a library in which the nucleic acid sequence immobilized on a substrate comprises a universal primer binding site (such as a PCR handle), a cell label (CL), a unique molecular identifier (UMI), a capture sequence (e.g., poly-T), and a target DNA sequence. Random primer extension can be performed to generate the library, and in some embodiments, the library can include P5 and P7 primer adapter sequences.

[0424] In some embodiments, compositions and methods include multiple constructions so that they are suitable for multiple labeling and / or analysis. For example, more than one binding reagent (e.g., protein binding reagent) can be provided, and the more than one binding reagent specifically binds to different target proteins and each includes different reagent oligonucleotides. In multiple constructions, more than one first oligonucleotide probe is provided, each having specificity for different reagent oligonucleotides. Eac...

Claims

1. A method for labeling DNA associated with a nuclear target in a cell, the method comprising: include: permeabilizing a cell comprising a nuclear target associated with double-stranded deoxyribonucleic acid (dsDNA), wherein optionally the dsDNA is genomic DNA (gDNA); contacting the nuclear target with a digestion composition comprising a DNA digestion enzyme and a binding agent capable of specifically binding to the nuclear target, wherein each binding agent comprises a binding agent-specific oligonucleotide comprising a unique identifier sequence for the binding agent, to produce more than one nuclear target-associated dsDNA fragments each comprising a single-stranded overhang; barcoding the more than one nuclear target associated dsDNA fragments or products thereof using first one or more oligonucleotide barcodes to generate more than one barcoded nuclear target associated DNA fragments, each of the more than one barcoded nuclear target associated DNA fragments comprising a sequence complementary to at least a portion of the nuclear target associated dsDNA fragments, wherein each of the first one or more oligonucleotide barcodes comprises a first target binding region capable of hybridizing to the more than one nuclear target associated dsDNA fragments or products thereof; and The binding reagent-specific oligonucleotide or its product is barcoded using a second one or more oligonucleotide barcode to produce more than one barcoded binding reagent-specific oligonucleotide, each of the more than one barcoded binding reagent-specific oligonucleotides comprising a sequence complementary to at least a portion of the unique identifier sequence, wherein each of the second one or more oligonucleotide barcodes comprises a second target binding region capable of hybridizing to the binding reagent-specific oligonucleotide or its product.

2. The method of claim 1, wherein contacting the nuclear target with the digestion composition comprises contacting the digestion composition with permeabilized cells.

3. The method of any one of claims 1-2, wherein contacting the nuclear target with the composition comprises the binding agent and the DNA digesting enzyme entering the permeabilized cell and the binding agent binding to the nuclear target.

4. The method of any one of claims 1-3, wherein the digestion composition comprises a fusion protein comprising the binding agent and the digestion enzyme, optionally wherein the fusion protein is sized to diffuse through the nuclear pores of the cell.

5. The method of any one of claims 1-3, wherein the digestive composition comprises a conjugate comprising the binding agent and the digestive enzyme, optionally wherein the conjugate is sized to diffuse through the nuclear pores of the cell.

6. The method according to any one of claims 1 to 3, wherein the DNA digesting enzyme comprises a domain capable of specifically binding to the binding agent, and optionally the domain of the DNA digesting enzyme comprises at least one of protein A, protein G, protein A / G or protein L. 7 . The method according to claim 6 , wherein the binding reagent and the DNA digesting enzyme are separated from each other when contacting with the permeabilized cells, and wherein the binding reagent and the DNA digesting enzyme enter the cells separately.

8. The method of any one of claims 6-7, wherein the DNA digesting enzyme binds to the binding agent within the nucleus of the cell.

9. The method of any one of claims 6-8, wherein the DNA digesting enzyme binds to the binding agent prior to entering the nucleus of the cell, and wherein the DNA digesting enzyme bound to the binding agent is sized to diffuse through the nuclear pores of the cell.

10. The method of any one of claims 6-9, wherein the DNA digesting enzyme binds to the binding agent before the binding agent binds to the nuclear target.

Citation Information

Patent Citations

  • Beet puller and topping machine.

    US1258367A

  • Digital Counting of Individual Molecules by Stochastic Attachment of Diverse Labels

    US20110160078A1

  • Massively parallel single cell analysis

    US20150299784A1

  • Measurement of protein expression using reagents with barcoded oligonucleotide sequences

    US20180088112A1

  • Synthetic multiplets for multiplets determination

    US20180276332A1