Preparation of Nucleic Acids for Further Analysis of Nucleic Acid Sequences

JP2025506418A5Pending Publication Date: 2026-02-13BECTON DICKINSON & CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024546189
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-02-07
Filing Date
2023-02-06
Publication Date
2026-02-13

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure includes systems, methods, compositions, and kits for full-length whole transcriptome analysis (WTA). Some embodiments include 5'-based, 3'-based, and internal-based gene expression profiling. In some embodiments, immune repertoire profiling methods are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Related Applications This application claims the benefit under 35 USC §119(e) of U.S. Provisional Patent Application No. 63 / 307,559, filed February 7, 2022, the contents of which are incorporated herein by reference in their entirety for all purposes. [Background technology]

[0002] The present disclosure relates generally to the field of molecular biology, and more specifically to multi-omics analysis using molecular barcoding. Molecular barcoding methods and techniques are useful for single cell transcriptomics analysis, such as deciphering gene expression profiles to determine the state of a cell using, for example, reverse transcription, polymerase chain reaction (PCR) amplification, and next generation sequencing (NGS). Molecular barcoding is also useful for single cell proteomics analysis. There is a need for methods and techniques for molecular barcoding at one or both of the 5' and 3' ends of a nucleic acid target. There is a need for compositions, systems, and methods that can efficiently quantitatively analyze gene expression of a cell. There is a need for compositions, systems, and methods that can obtain the sequence of an internal region of a nucleic acid target. There is a need for compositions, systems, and methods that can obtain the full length sequence of a nucleic acid target. Summary of the Invention

[0003] The disclosure herein includes a method for labeling a nucleic acid target in a sample. In some embodiments, the method includes contacting a copy of the nucleic acid target with a first plurality of oligonucleotide barcodes, each oligonucleotide barcode of the first plurality of oligonucleotide barcodes comprising a first universal sequence, a first molecular label, and a first target binding region capable of hybridizing to the nucleic acid target. In some embodiments, the method includes extending the first plurality of oligonucleotide barcodes hybridized to the copy of the nucleic acid target to generate a plurality of barcoded nucleic acid molecules, each of which comprises the first universal sequence, the first molecular label, and a sequence complementary to at least a portion of the nucleic acid target. In some embodiments, the method includes contacting the barcoded nucleic acid molecule with a second plurality of oligonucleotide barcodes for hybridization. In some embodiments, each oligonucleotide barcode of the second plurality of oligonucleotide barcodes comprises a second universal sequence, a cleavage domain, and a blocking group. In some embodiments, a blocking group can prevent extension of the oligonucleotide barcode and a cleavage domain is located 5' to the blocking group, such that when the cleavage domain hybridizes to the barcoded nucleic acid molecule, the oligonucleotide barcode can be cleaved by a cleavage enzyme at a point within or adjacent to the cleavage domain. In some embodiments, the method includes contacting a second plurality of oligonucleotide barcodes hybridized to the barcoded nucleic acid molecule with a cleavage enzyme, thereby removing the blocking group from the oligonucleotide barcodes. In some embodiments, the method includes extending a 3' end of an oligonucleotide barcode of the second plurality of oligonucleotide barcodes hybridized to the barcoded nucleic acid molecule to generate a plurality of extended barcoded nucleic acid molecules.

[0004] In some embodiments, each extended barcoded nucleic acid molecule of the plurality of extended barcoded nucleic acid molecules comprises at least a partial sequence of a nucleic acid target. In some embodiments, the cleavage enzyme is a ribonuclease H enzyme and / or a ribonuclease H2 enzyme, and the ribonuclease H2 enzyme may be a Pyrococcus abyssi ribonuclease H2 enzyme. In some embodiments, the cleavage enzyme is a hot-start cleavage enzyme that is thermostable and has reduced activity at low temperatures. In some embodiments, the hot-start cleavage enzyme is a Pyrococcus abyssi ribonuclease H2 that includes (a) a G12A amino acid substitution; (b) a P13T amino acid substitution; (c) a G169A amino acid substitution; or (d) a combination thereof. In some embodiments, the cleavage enzyme is chemically modified. In some embodiments, the cleavage enzyme is a chemically modified hot-start cleavage enzyme that is thermostable and has reduced activity at low temperatures, and the cleavage enzyme may be reversibly inactivated by interaction with an antibody at low temperatures. In some embodiments, the cleavage domain comprises one or more ribonucleotides capable of being cleaved by a RNase H enzyme. In some embodiments, the cleavage domain comprises one or more of the following moieties: a DNA residue, an abasic residue, a modified nucleoside, or a modified phosphate internucleotide linkage. In some embodiments, the cleavage domain comprises at least one RNA base. In some embodiments, the cleavage domain comprises one or more 2'-modified nucleosides, and the one or more modified nucleosides may be 2'-fluoro nucleosides. In some embodiments, the blocking group is attached to the 3' terminal nucleotide of the oligonucleotide barcode. In some embodiments, the blocking group is at or near the 3' end of the oligonucleotide barcode. In some embodiments, the blocking group is a 2',3'-dideoxynucleotide, a ribonucleotide residue, a 2',3'-SH nucleotide, or a 2'-O-PO3 nucleotide. In some embodiments, the blocking group comprises a non-nucleotide modification. In some embodiments, the blocking group further comprises a naphthyl-azo compound, a spacer, and / or biotin.

[0005] In some embodiments, extending the 3' ends of the oligonucleotide barcodes of the second plurality of oligonucleotide barcodes comprises extending the 3' ends of the oligonucleotide barcodes of the second plurality of oligonucleotide barcodes with a DNA polymerase having strand displacement activity. In some embodiments, extending the 3' ends of the oligonucleotide barcodes of the second plurality of oligonucleotide barcodes with a DNA polymerase having strand displacement activity can generate an extended barcoded nucleic acid molecule comprising a complement of the first molecular label and a complement of the first universal sequence. In some embodiments, extending the 3' ends of the oligonucleotide barcodes of the second plurality of oligonucleotide barcodes comprises extending the 3' ends of the oligonucleotide barcodes of the second plurality of oligonucleotide barcodes with a DNA polymerase not having strand displacement activity.In some embodiments, the polymerase is selected from the group consisting of Phi29 DNA polymerase, E. coli DNA polymerase I, Bsu DNA polymerase, Bst DNA polymerase, Taq DNA polymerase, VENT™ DNA polymerase, DEEPVENT™ DNA polymerase, LongAmp® Taq DNA polymerase, LongAmp® Hot Start Taq DNA polymerase, Crimson LongAmp® Taq DNA polymerase, Crimson Taq DNA polymerase, OneTaq® DNA polymerase, OneTaq® Quick-Load® DNA polymerase, Hemo KlenTaq® DNA polymerase, REDTaq® DNA polymerase, Phusion® DNA polymerase, Phusion® High-Fidelity DNA polymerase, Platinum Pfx DNA polymerase, AccuPrime Pfx DNA polymerase, Klenow fragment, Pwo DNA polymerase, Pfu The 3' end of the oligonucleotide barcode is selected from the group consisting of a DNA polymerase, a T4 DNA polymerase, a T7 DNA polymerase, derivatives thereof, or any combination thereof. In some embodiments, extending the 3' end of the oligonucleotide barcode comprises extending the 3' end of the oligonucleotide barcode using a mesophilic DNA polymerase, a thermophilic DNA polymerase, a psychrophilic DNA polymerase, or any combination thereof. In some embodiments, extending the 3' end of the oligonucleotide barcode comprises extending the 3' end of the oligonucleotide barcode using a DNA polymerase lacking at least one of 5' to 3' exonuclease activity and 3' to 5' exonuclease activity, wherein the DNA polymerase may comprise a Klenow fragment. In some embodiments, extending the first plurality of oligonucleotide barcodes comprises extending the first plurality of oligonucleotide barcodes using a reverse transcriptase. In some embodiments, the reverse transcriptase is capable of terminal transferase activity.In some embodiments, the reverse transcriptase with strand displacement activity is PrimeScript reverse transcriptase, M-MuLV reverse transcriptase, SmartScribe reverse transcriptase, Maxima H Minus reverse transcriptase, and / or Superscript II reverse transcriptase. In some embodiments, the reverse transcriptase comprises a viral reverse transcriptase, which may be murine leukemia virus (MLV) reverse transcriptase or Moloney murine leukemia virus (MMLV) reverse transcriptase.

[0006] In some embodiments, each oligonucleotide barcode of the second plurality of oligonucleotide barcodes comprises a second molecular label, and at least 10 of the second plurality of oligonucleotide barcodes comprise a different second molecular label sequence, and each second molecular label may comprise at least 6 nucleotides, and further, the second molecular label sequence may be a random sequence. In some embodiments, the second plurality of oligonucleotide barcodes hybridize to the barcoded nucleic acid molecule by hybridization between the second molecular label and a sequence complementary to at least a portion of the nucleic acid target. In some embodiments, each oligonucleotide barcode of the second plurality of oligonucleotide barcodes comprises a second target binding region. In some embodiments, the first target binding region and / or the second target binding region comprises a poly(dA) region, a poly(dT) region, a random sequence, a gene-specific sequence, or any combination thereof. In some embodiments, the oligonucleotide barcodes of the second plurality of oligonucleotide barcodes hybridize to the barcoded nucleic acid molecule by hybridization between the second target binding region and a sequence complementary to at least a portion of the nucleic acid target. In some embodiments, at least ten of the second plurality of oligonucleotide barcodes comprise different second target binding regions, and at least two of the target binding regions may be capable of binding to the complement of different nucleic acid targets, and further, at least two of the target binding regions may be capable of hybridizing to different regions of the complement of the same nucleic acid target. In some embodiments, two or more oligonucleotide barcodes of the second plurality of oligonucleotide barcodes may hybridize to different regions of the complement of the same nucleic acid target to generate two or more extended barcoded nucleic acid molecules. In some embodiments, the two or more extended barcoded nucleic acid molecules may be generated by two or more oligonucleotide barcodes of the second plurality of oligonucleotide barcodes hybridizing to different regions of the complement of the same nucleic acid target.In some embodiments, the two or more extended barcoded nucleic acid molecules collectively comprise at least about 50% of the entire sequence of the nucleic acid target.

[0007] In some embodiments, the method includes denaturing a plurality of barcoded nucleic acid molecules. In some embodiments, the method includes denaturing a plurality of extended barcoded nucleic acid molecules. In some embodiments, the method includes determining the copy number of a nucleic acid target in a sample based on the number of first molecular labels having distinct sequences associated with the plurality of barcoded nucleic acid molecules or products thereof. In some embodiments, the method includes determining the copy number of a nucleic acid target in a sample based on the number of first molecular labels having distinct sequences, second molecular labels having distinct sequences, or a combination thereof associated with the plurality of extended barcoded nucleic acid molecules or products thereof. In some embodiments, determining the copy number of the nucleic acid target comprises determining the copy number of each of the plurality of nucleic acid targets in the sample based on the number of first molecular labels with distinct sequences associated with a barcoded nucleic acid molecule of the plurality of barcoded nucleic acid molecules, or a product thereof, that comprises the sequence of each of the plurality of nucleic acid targets; and / or the number of first molecular labels with distinct sequences, second molecular labels with distinct sequences, or a combination thereof, associated with an extended barcoded nucleic acid molecule of the plurality of extended barcoded nucleic acid molecules, that comprises the sequence of each of the plurality of nucleic acid targets. In some embodiments, the sequence of each of the plurality of nucleic acid targets comprises a subsequence of each of the plurality of nucleic acid targets. In some embodiments, the sequence of a nucleic acid target within the plurality of barcoded nucleic acid molecules comprises a subsequence of a nucleic acid target. In some embodiments, the nucleic acid target comprises mRNA. In some embodiments, the sample comprises a single cell, optionally an immune cell, and further optionally a B cell or a T cell. In some embodiments, the sample comprises a plurality of cells, a plurality of single cells, a tissue, a tumor sample, or any combination thereof. In some embodiments, the single cell comprises a circulating tumor cell.

[0008] In some embodiments, the first universal sequence of each oligonucleotide barcode of the first plurality of oligonucleotide barcodes is 5' to the first molecular label and the first target binding region; and / or the second universal sequence of each oligonucleotide barcode of the second plurality of oligonucleotide barcodes is 5' to the second molecular label and / or the second target binding region.

[0009] In some embodiments, the method includes amplifying a plurality of barcoded nucleic acid molecules using an amplification primer and a primer comprising a first universal sequence or a portion thereof, thereby generating a first plurality of single-labeled nucleic acid molecules comprising a sequence or a portion thereof of the nucleic acid target, and determining the copy number of the nucleic acid target in the sample includes determining the copy number of the nucleic acid target in the sample based on the number of first molecular labels having distinct sequences associated with the first plurality of single-labeled nucleic acid molecules or products thereof. In some embodiments, the method includes amplifying a plurality of extended barcoded nucleic acid molecules using an amplification primer and a primer comprising a second universal sequence or a portion thereof, thereby generating a second plurality of single-labeled nucleic acid molecules comprising a sequence or a portion thereof of the nucleic acid target, and determining the copy number of the nucleic acid target in the sample includes determining the copy number of the nucleic acid target in the sample based on the number of second molecular labels having distinct sequences associated with the second plurality of single-labeled nucleic acid molecules or products thereof. In some embodiments, the amplification primer comprises a fourth universal sequence. In some embodiments, the amplification primer is a target-specific primer. In some embodiments, the target-specific primers specifically hybridize to an immune receptor, a constant region of an immune receptor, a variable region of an immune receptor, a diversity region of an immune receptor, and / or a junction of a variable region and a diversity region of an immune receptor. In some embodiments, the immune receptor is a T cell receptor (TCR) and / or a B cell receptor (BCR) receptor, where the TCR may comprise a TCR alpha chain, a TCR beta chain, a TCR gamma chain, a TCR delta chain, or any combination thereof, and the BCR receptor comprises a BCR heavy chain and / or a BCR light chain.

[0010] In some embodiments, the method includes hybridizing random primers to a plurality of barcoded nucleic acid molecules and extending the random primers to generate a first plurality of extension products, where the random primers comprise a third universal sequence or its complement; and amplifying the first plurality of extension products using a primer capable of hybridizing to the third universal sequence or its complement and a primer capable of hybridizing to the first universal sequence or its complement, thereby generating a third plurality of single-labeled nucleic acid molecules or products thereof. In some embodiments, determining the copy number of the nucleic acid target in the sample includes determining the copy number of the nucleic acid target in the sample based on the number of first molecular labels having distinct sequences associated with the third plurality of single-labeled nucleic acid molecules or products thereof. In some embodiments, the method includes hybridizing a random primer to the plurality of extended barcoded nucleic acid molecules and extending the random primer to generate a second plurality of extension products, where the random primer comprises a third universal sequence or its complement; and amplifying the second plurality of extension products using a primer capable of hybridizing to the third universal sequence or its complement and a primer capable of hybridizing to the second universal sequence or its complement, thereby generating a fourth plurality of single-labeled nucleic acid molecules or products thereof. In some embodiments, determining the copy number of the nucleic acid target in the sample includes determining the copy number of the nucleic acid target in the sample based on the number of second molecular labels having distinct sequences associated with the fourth plurality of single-labeled nucleic acid molecules or products thereof.

[0011] In some embodiments, the first universal sequence, the second universal sequence, the third universal sequence, and / or the fourth universal sequence are the same. In some embodiments, the first universal sequence, the second universal sequence, the third universal sequence, and / or the fourth universal sequence are different. In some embodiments, the first universal sequence, the second universal sequence, the third universal sequence, and / or the fourth universal sequence comprise the binding site of the sequencing primer and / or the sequencing adaptor, their complementary sequence, and / or a portion thereof. In some embodiments, the sequencing adaptor comprises a P5 sequence, a P7 sequence, their complementary sequence, and / or a portion thereof. In some embodiments, the sequencing primer comprises a lead 1 sequencing primer, a lead 2 sequencing primer, their complementary sequence, and / or a portion thereof.

[0012] In some embodiments, the method includes obtaining sequence information of a plurality of barcoded nucleic acid molecules or products thereof. In some embodiments, obtaining sequence information includes attaching a sequencing adapter to a plurality of extended barcoded nucleic acid molecules or products thereof. In some embodiments, the method includes obtaining sequence information of a plurality of extended barcoded nucleic acid molecules or products thereof. In some embodiments, obtaining sequence information includes attaching a sequencing adapter to an extended barcoded nucleic acid molecule, a barcoded nucleic acid molecule, a product thereof, or any combination thereof. In some embodiments, the method includes obtaining sequence information of one or more of the first, second, third, and fourth plurality of single-labeled nucleic acid molecules or products thereof. In some embodiments, obtaining sequence information includes attaching a sequencing adapter to one or more of the first, second, third, and fourth plurality of single-labeled nucleic acid molecules or products thereof. In some embodiments, obtaining sequence information of one or more of the first, second, third, and fourth plurality of single-labeled nucleic acid molecules or products thereof comprises obtaining sequencing data comprising a plurality of sequencing reads of one or more of the first, second, third, and fourth plurality of single-labeled nucleic acid molecules or products thereof, each of the plurality of sequencing reads comprising (1) a cellular label sequence, (2) a molecular label sequence, and / or (3) a subsequence of the nucleic acid target.

[0013] In some embodiments, the method includes aligning each of a plurality of sequencing reads of the nucleic acid target with each unique cell label sequence indicative of a single cell of the sample to generate an aligned sequence of the nucleic acid target. In some embodiments, the aligned sequence of the nucleic acid target includes at least 50% of the cDNA sequence of the nucleic acid target, at least 70% of the cDNA sequence of the nucleic acid target, at least 90% of the cDNA sequence of the nucleic acid target, or the full length of the cDNA sequence of the nucleic acid target. In some embodiments, the nucleic acid target is an immune receptor, and the immune receptor may include a BCR light chain, a BCR heavy chain, a TCR alpha chain, a TCR beta chain, a TCR gamma chain, a TCR delta chain, or any combination thereof. In some embodiments, the aligned sequence of the nucleic acid target includes a complementarity determining region 1 (CDR1), a complementarity determining region 2 (CDR2), a complementarity determining region 3 (CDR3), a variable region, a full length of a variable region, or a combination thereof. In some embodiments, the aligned sequences of the nucleic acid targets include variable regions, diversity regions, junctions of variable regions, diversity regions and / or constant regions, or any combination thereof. In some embodiments, obtaining sequence information includes obtaining sequence information of the BCR light chain and BCR heavy chain of the single cell, and the sequence information of the BCR light chain and BCR heavy chain may include the sequence of the complementarity determining region 1 (CDR1), CDR2, CDR3, or any combination thereof, of the BCR light chain and / or BCR heavy chain. In some embodiments, the method includes pairing the BCR light chain and BCR heavy chain of the single cell based on the obtained sequence information. In some embodiments, the sample includes a plurality of single cells, and the method includes pairing the BCR light chain and BCR heavy chain of at least 50% of the single cells based on the obtained sequence information. In some embodiments, obtaining sequence information comprises obtaining sequence information of the TCR alpha chain and the TCR beta chain of the single cell, and the sequence information of the TCR alpha chain and the TCR beta chain may comprise the sequence of complementarity determining region 1 (CDR1), CDR2, CDR3, or any combination thereof, of the TCR alpha chain and / or the TCR beta chain. In some embodiments, the method comprises pairing the TCR alpha chain and the TCR beta chain of the single cell based on the obtained sequence information.In some embodiments, the sample comprises a plurality of single cells, and the method comprises pairing the TCR alpha chain and the TCR beta chain of at least 50% of the single cells based on the sequence information obtained. In some embodiments, obtaining sequence information comprises obtaining sequence information of the TCR gamma chain and the TCR delta chain of the single cells. In some embodiments, the sequence information of the TCR gamma chain and the TCR delta chain comprises the sequence of the complementarity determining region 1 (CDR1), CDR2, CDR3, or any combination thereof, of the TCR gamma chain and / or the TCR delta chain. In some embodiments, the method comprises pairing the TCR gamma chain and the TCR delta chain of the single cells based on the sequence information obtained. In some embodiments, the sample comprises a plurality of single cells, and the method comprises pairing the TCR gamma chain and the TCR delta chain of at least 50% of the single cells based on the sequence information obtained.

[0014] In some embodiments, the complement of the molecular label comprises a reverse complement sequence of the molecular label or a complementary sequence of the molecular label. In some embodiments, the plurality of barcoded nucleic acid molecules comprises barcoded deoxyribonucleic acid (DNA) molecules, barcoded ribonucleic acid (RNA) molecules, or a combination thereof. In some embodiments, the nucleic acid target comprises a nucleic acid molecule, the nucleic acid molecule may comprise ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNA comprising a poly(A) tail, or any combination thereof, and further, the mRNA may encode an immune receptor. In some embodiments, the nucleic acid target comprises a cellular component binding reagent, and / or the nucleic acid molecule is associated with a cellular component binding reagent, and the method may further comprise dissociating the nucleic acid molecule and the cellular component binding reagent. In some embodiments, at least 10 of the first and / or second plurality of oligonucleotide barcodes comprise different molecular label sequences. In some embodiments, each molecular label of the first and / or second plurality of oligonucleotide barcodes comprises at least 6 nucleotides. In some embodiments, the first and / or second plurality of oligonucleotide barcodes are associated with a solid support. In some embodiments, the first and / or second plurality of oligonucleotide barcodes associated with the same solid support each comprise the same sample label. In some embodiments, each sample label of the first and / or second plurality of oligonucleotide barcodes comprises at least 6 nucleotides. In some embodiments, the first and / or second plurality of oligonucleotide barcodes each comprise a cell label. In some embodiments, each cell label of the first and / or second plurality of oligonucleotide barcodes comprises at least 6 nucleotides. In some embodiments, the oligonucleotide barcodes of the first and / or second plurality of oligonucleotide barcodes associated with the same solid support comprise the same cell label. In some embodiments, the oligonucleotide barcodes of the first and / or second plurality of oligonucleotide barcodes associated with different solid supports comprise different cell labels.In some embodiments, the method includes extending the oligonucleotide barcode in the presence of one or more of ethylene glycol, polyethylene glycol, 1,2-propanediol, dimethylsulfoxide (DMSO), glycerol, formamide, 7-deaza-GTP, acetamide, tetramethylammonium chloride salts, betaine, or any combination thereof.

[0015] In some embodiments, the solid support comprises a synthetic particle, a planar surface, or a combination thereof. In some embodiments, the sample comprises a single cell, and the method comprises associating a synthetic particle comprising a first and a second plurality of oligonucleotide barcodes with the single cell in the sample. In some embodiments, the method comprises lysing the single cell after associating the synthetic particle with the single cell, where lysing the single cell may comprise heating the sample, contacting the sample with a detergent, changing the pH of the sample, or any combination thereof. In some embodiments, the synthetic particle and the single cell are in the same compartment, which may be a well or a droplet. In some embodiments, at least one oligonucleotide barcode of the first and / or second plurality of oligonucleotide barcodes is immobilized or partially immobilized on the synthetic particle, or at least one oligonucleotide barcode of the first and / or second plurality of oligonucleotide barcodes is encapsulated or partially encapsulated within the synthetic particle. In some embodiments, the synthetic particle is disintegrable, and may be a disintegrable hydrogel particle. In some embodiments, the synthetic particles comprise beads, which may be sepharose beads, streptavidin beads, agarose beads, magnetic beads, conjugated beads, protein A conjugated beads, protein G conjugated beads, protein A / G conjugated beads, protein L conjugated beads, oligo(dT) conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, or any combination thereof. In some embodiments, the synthetic particles comprise a material selected from the group consisting of polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic, ceramic, plastic, glass, methylstyrene, acrylic polymer, titanium, latex, sepharose, cellulose, nylon, silicone, or any combination thereof.In some embodiments, each oligonucleotide barcode of the first and / or second plurality of oligonucleotide barcodes comprises a linker functional group. In some embodiments, the synthetic particle comprises a solid support functional group. In some embodiments, the support functional group and the linker functional group are associated with each other, and the linker functional group and the support functional group may each be selected from the group consisting of C6, biotin, streptavidin, primary amines, aldehydes, ketones, and any combination thereof.

[0016] In some embodiments, a solid support is provided. Disclosed herein in some embodiments is a solid support associated with one or both of a first and second plurality of oligonucleotide barcodes. [Brief description of the drawings]

[0017] [Figure 1] FIG. 1 illustrates a non-limiting exemplary barcode. [Diagram 2] FIG. 1 illustrates a non-limiting exemplary workflow of barcoding and electronic counting. [Diagram 3] 1 is a schematic diagram showing a non-limiting exemplary process for generating an indexed library of 3′-barcoded targets from multiple targets. [Figure 4A] 1 shows a schematic diagram of a non-limiting exemplary workflow for determining the sequence of a nucleic acid target (e.g., the V(D)J region of an immune receptor) and full-length whole transcriptome analysis (WTA) using the barcoding methods and compositions provided herein. [Figure 4B] 1 shows a schematic diagram of a non-limiting exemplary workflow for determining the sequence of a nucleic acid target (e.g., the V(D)J region of an immune receptor) and full-length whole transcriptome analysis (WTA) using the barcoding methods and compositions provided herein. [Figure 4C]1 shows a schematic diagram of a non-limiting exemplary workflow for determining the sequence of a nucleic acid target (e.g., the V(D)J region of an immune receptor) and full-length whole transcriptome analysis (WTA) using the barcoding methods and compositions provided herein. [Figure 4D] 1 shows a schematic diagram of a non-limiting exemplary workflow for determining the sequence of a nucleic acid target (e.g., the V(D)J region of an immune receptor) and full-length whole transcriptome analysis (WTA) using the barcoding methods and compositions provided herein. [Figure 4E] 1 shows a schematic diagram of a non-limiting exemplary workflow for determining the sequence of a nucleic acid target (e.g., the V(D)J region of an immune receptor) and full-length whole transcriptome analysis (WTA) using the barcoding methods and compositions provided herein. [Figure 4F] 1 shows a schematic diagram of a non-limiting exemplary workflow for determining the sequence of a nucleic acid target (e.g., the V(D)J region of an immune receptor) and full-length whole transcriptome analysis (WTA) using the barcoding methods and compositions provided herein. [Figure 4G] 1 shows a schematic diagram of a non-limiting exemplary workflow for determining the sequence of a nucleic acid target (e.g., the V(D)J region of an immune receptor) and full-length whole transcriptome analysis (WTA) using the barcoding methods and compositions provided herein. [Figure 4H] 1 shows a schematic diagram of a non-limiting exemplary workflow for determining the sequence of a nucleic acid target (e.g., the V(D)J region of an immune receptor) and full-length whole transcriptome analysis (WTA) using the barcoding methods and compositions provided herein. [Figure 5A] 1 shows a schematic diagram of a non-limiting exemplary workflow for determining the sequence of a nucleic acid target (e.g., the V(D)J region of an immune receptor) and full-length whole transcriptome analysis (WTA) using the barcoding methods and compositions provided herein. [Figure 5B]1 shows a schematic diagram of a non-limiting exemplary workflow for determining the sequence of a nucleic acid target (e.g., the V(D)J region of an immune receptor) and full-length whole transcriptome analysis (WTA) using the barcoding methods and compositions provided herein. [Figure 5C] 1 shows a schematic diagram of a non-limiting exemplary workflow for determining the sequence of a nucleic acid target (e.g., the V(D)J region of an immune receptor) and full-length whole transcriptome analysis (WTA) using the barcoding methods and compositions provided herein. [Figure 5D] 1 shows a schematic diagram of a non-limiting exemplary workflow for determining the sequence of a nucleic acid target (e.g., the V(D)J region of an immune receptor) and full-length whole transcriptome analysis (WTA) using the barcoding methods and compositions provided herein. [Figure 5E] 1 shows a schematic diagram of a non-limiting exemplary workflow for determining the sequence of a nucleic acid target (e.g., the V(D)J region of an immune receptor) and full-length whole transcriptome analysis (WTA) using the barcoding methods and compositions provided herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0018] In the following detailed description, reference is made to the accompanying drawings, which form a part of this specification. In the drawings, similar symbols typically identify similar components unless otherwise indicated by context. The exemplary embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments can be used and other changes can be made without departing from the spirit or scope of the subject matter presented herein. It is readily understood that the aspects of the present disclosure, as generally described herein and illustrated in the drawings, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are expressly contemplated herein and form part of the disclosure of this specification. All patents, published patent applications, other publications, and sequences from GenBank and other databases referenced herein are incorporated by reference in their entirety for relevant art.

[0019] Quantifying a small number of nucleic acids, such as messenger ribonucleotide acid (mRNA) molecules, is clinically important, for example, to determine genes expressed in cells at different developmental stages or under different environmental conditions. However, it can also be very difficult to determine the absolute number of nucleic acid molecules (e.g., mRNA molecules), especially when the number of molecules is very small. One method to determine the absolute number of molecules in a sample is digital polymerase chain reaction (PCR). Ideally, PCR produces identical molecular copies at each cycle. However, PCR can have drawbacks because each molecule replicates with a stochastic probability, which varies with PCR cycles and gene sequence, resulting in amplification bias and inaccurate gene expression measurements. Stochastic barcodes with unique molecular labels (also called molecular indexes (MI)) can be used to count the number of molecules and correct for amplification bias. Stochastic barcoding, such as the Precise™ assay (Cellular Research, Inc., Palo Alto, Calif.) and the Rhapsody™ assay (Becton, Dickinson and Company, Franklin Lakes, NJ), can correct for biases induced by the library preparation step by using molecular beacons (MLs) to label mRNA during PCR and reverse transcription (RT).

[0020] The Precise™ assay utilizes a non-exhaustive pool of stochastic barcodes with a large number, e.g., 6561-65536, of unique molecular label sequences on poly(T) oligonucleotides to hybridize to all poly(A)-mRNAs in a sample during the RT step. The stochastic barcodes may contain universal PCR priming sites. During RT, target gene molecules react randomly with the stochastic barcodes. Each target molecule may hybridize to a stochastic barcode, thereby generating a stochastically barcoded complementary ribonucleotide acid (cDNA) molecule. After labeling, the stochastically barcoded cDNA molecules from the microwells of the microwell plate may be pooled into a single tube for PCR amplification and sequencing. The raw sequencing data may be analyzed to obtain the number of reads, the number of stochastic barcodes with unique molecular label sequences, and the number of mRNA molecules.

[0021] The disclosure herein includes a method for labeling a nucleic acid target in a sample. In some embodiments, the method includes contacting a copy of the nucleic acid target with a first plurality of oligonucleotide barcodes, each oligonucleotide barcode of the first plurality of oligonucleotide barcodes comprising a first universal sequence, a first molecular label, and a first target binding region capable of hybridizing to the nucleic acid target. In some embodiments, the method includes extending the first plurality of oligonucleotide barcodes hybridized to the copy of the nucleic acid target to generate a plurality of barcoded nucleic acid molecules, each of which comprises the first universal sequence, the first molecular label, and a sequence complementary to at least a portion of the nucleic acid target. In some embodiments, the method includes contacting the barcoded nucleic acid molecule with a second plurality of oligonucleotide barcodes for hybridization. In some embodiments, each oligonucleotide barcode of the second plurality of oligonucleotide barcodes comprises a second universal sequence, a cleavage domain, and a blocking group. In some embodiments, a blocking group can prevent extension of the oligonucleotide barcode, and a cleavage domain is located 5' to the blocking group, such that when the cleavage domain hybridizes to the barcoded nucleic acid molecule, the oligonucleotide barcode can be cleaved by a cleavage enzyme at a point within or adjacent to the cleavage domain. In some embodiments, the method includes contacting a second plurality of oligonucleotide barcodes hybridized to the barcoded nucleic acid molecule with a cleavage enzyme, thereby removing the blocking group from the oligonucleotide barcodes. In some embodiments, the method includes extending a 3' end of an oligonucleotide barcode of the second plurality of oligonucleotide barcodes hybridized to the barcoded nucleic acid molecule to generate a plurality of extended barcoded nucleic acid molecules. In some embodiments, a solid support is provided. In some embodiments, a solid support associated with one or both of the first and second plurality of oligonucleotide barcodes is disclosed herein.

[0022] definition Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this disclosure belongs. See, for example, Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989). For purposes of this disclosure, the following terms are defined below.

[0023] As used herein, the term "adapter" may refer to a sequence for facilitating amplification or sequencing of an associated nucleic acid. The associated nucleic acid may include a target nucleic acid. The associated nucleic acid may include one or more of a spatial label, a target label, a sample label, an indexing label, or a barcode sequence (e.g., a molecular label). The adapter may be linear. The adapter may be a pre-adenylated adapter. The adapter may be double-stranded or single-stranded. One or more adapters may be placed at the 5' or 3' end of the nucleic acid. When the adapter includes a known sequence at the 5' and 3' ends, the known sequence may be the same or different sequences. The adapters placed at the 5' and / or 3' ends of the polynucleotide may be capable of hybridizing to one or more oligonucleotides immobilized on a surface. The adapter may include a universal sequence in some embodiments. The universal sequence may be a region of nucleotide sequence that is common to two or more nucleic acid molecules. The two or more nucleic acid molecules may also have regions of different sequences. Thus, for example, the 5' adaptors may comprise the same and / or universal nucleic acid sequences and the 3' adaptors may comprise the same and / or universal sequences. The universal sequences may be present in different members of a plurality of nucleic acid molecules, thereby allowing the replication or amplification of multiple different sequences using a single universal primer that is complementary to the universal sequence. Similarly, at least one, two (e.g., a pair) or more universal sequences may be present in different members of a collection of nucleic acid molecules, thereby allowing the replication or amplification of multiple different sequences using at least one, two (e.g., a pair) or more single universal primers that are complementary to the universal sequence. Thus, a universal primer includes a sequence that can hybridize to such a universal sequence. A molecule having a target nucleic acid sequence may be modified to add a universal adaptor (e.g., a non-target nucleic acid sequence) to one or both ends of the different target nucleic acid sequences.The one or more universal primers bound to the target nucleic acid may provide a site for hybridization of the universal primer. The one or more universal primers bound to the target nucleic acid may be the same or different from each other.

[0024] As used herein, the term "associated" or "associated with" may mean that two or more species are identifiable as being located together at a time. Association may mean that two or more species are or were in similar containers. Association may also be an informational association. For example, digital information about two or more species may be stored and used to determine that one or more of the species were located together at a time. Association may also be a physical association. In some embodiments, two or more associated species are "tethered," "bound," or "immobilized" to each other or to a common solid or semi-solid surface. Association may refer to a covalent or non-covalent means for attaching a label to a solid or semi-solid support such as a bead. Association may be a covalent bond between a target and a label. Association may include hybridization between two molecules (e.g., a target molecule and a label).

[0025] As used herein, the term "complementary" may refer to the ability for exact pairing between two nucleotides. For example, if a nucleotide at a given position of a nucleic acid can hydrogen bond with a nucleotide of another nucleic acid, the two nucleic acids are considered to be complementary to each other at that position. Complementarity between two single-stranded nucleic acid molecules may be "partial," where only some of the nucleotides bind, or may be complete, where there is total complementarity between the single-stranded molecules. A first nucleotide sequence may be referred to as the "complement" of a second sequence if the first nucleotide sequence is complementary to the second nucleotide sequence. A first nucleotide sequence may be referred to as the "reverse complement" of a second sequence if the first nucleotide sequence is complementary to a sequence that is the reverse of the second sequence (i.e., the order of the nucleotides is reversed). As used herein, a "complementary" sequence may refer to the "complement" or "reverse complement" of a sequence. It is understood from this disclosure that when a molecule is capable of hybridizing to another molecule, it may be complementary or partially complementary to the hybridizing molecule.

[0026] As used herein, the term "digital counting" can refer to a method for estimating the number of target molecules in a sample. Digital counting can include determining the number of unique labels associated with targets in a sample. This methodology, which can be probabilistic in nature, converts the problem of molecular counting into one of locating and identifying identical molecules into a series of yes / no digital questions regarding the detection of a predefined set of labels.

[0027] As used herein, the term "label" or "labels" may refer to a nucleic acid code associated with a target in a sample. The label may be, for example, a nucleic acid label. The label may be a fully or partially amplifiable label. The label may be a fully or partially sequenceable label. The label may be a portion of a naturally occurring nucleic acid that can be identified as distinct. The label may be a known sequence. The label may include a junction of a nucleic acid sequence, for example, a junction of a naturally occurring and a non-natural sequence. As used herein, the term "label" may be used interchangeably with the terms "index," "tag," or "label tag." The label may carry information. For example, in various embodiments, the label may be used to determine the identity of the sample, the source of the sample, the identity of the cell, and / or the target.

[0028] As used herein, the term "non-depleting reservoir" may refer to a pool of barcodes (e.g., stochastic barcodes) that are composed of a number of different labels. A non-depleting reservoir may contain a number of different barcodes such that when the non-depleting reservoir is associated with a pool of targets, each target is more likely to be associated with a unique barcode. The uniqueness of each labeled target molecule can be determined by random selection statistics and depends on the number of copies of the same target molecule in the population compared to the diversity of the labels. The size of the resulting labeled target molecule can be determined by the stochastic nature of the barcoding process, and analysis of the number of barcodes detected then allows for the calculation of the number of target molecules present in the original population or sample. If the ratio of the number of target molecules present to the number of unique barcodes is low, the labeled target molecule is highly unique (i.e., the probability that more than one target molecule will be labeled with one given label is very low).

[0029] As used herein, the term "nucleic acid" refers to a polynucleotide sequence or a fragment thereof. A nucleic acid may comprise nucleotides. A nucleic acid may be exogenous or endogenous to a cell. A nucleic acid may be present in a cell-free environment. A nucleic acid may be a gene or a fragment thereof. A nucleic acid may be DNA. A nucleic acid may be RNA. A nucleic acid may comprise one or more analogs (e.g., modified backbones, sugars, or nucleobases). Some non-limiting examples of analogs include 5-bromouracil, peptide nucleic acid, xenonucleic acid, morpholino, locked nucleic acid, glycol nucleic acid, threose nucleic acid, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein linked to a sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queosine, and wyosine. "Nucleic acid," "polynucleotide," "target polynucleotide," and "target nucleic acid" can be used interchangeably.

[0030] The nucleic acid may include one or more modifications (e.g., base modifications, backbone modifications) to result in a nucleic acid with new or enhanced properties (e.g., improved stability). The nucleic acid may include a nucleic acid affinity tag. A nucleoside may be a base-sugar combination. The base portion of the nucleoside may be a heterocyclic base. The two most common classes of such heterocyclic bases are purines and pyrimidines. A nucleotide may be a nucleoside further comprising a phosphate group covalently linked to the sugar portion of the nucleoside. For nucleosides that include a pentofuranosyl sugar, the phosphate group may be linked to the 2', 3', or 5' hydroxyl moiety of the sugar. In forming a nucleic acid, the phosphate group may covalently link adjacent nucleosides to each other to form a linear polymeric compound. The respective ends of this linear polymeric compound may then be further joined to form a circular compound, although linear compounds are generally preferred. In addition, linear compounds may have internal nucleotide base complementarity and therefore fold in such a manner that a fully or partially double-stranded compound results. Within nucleic acids, the phosphate groups may be commonly referred to as forming the internucleoside backbone of the nucleic acid. The linkage or backbone may be a 3' and 5' phosphodiester linkage.

[0031] The nucleic acids may contain modified backbones and / or modified internucleoside linkages. Modified backbones can include those that retain a phosphorus atom in the backbone and those that do not have a phosphorus atom in the backbone. Suitable modified nucleic acid backbones containing a phosphorus atom therein can include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkyl phosphotriesters, methyl and other alkyl phosphonates, such as 3'-alkylene phosphonates, 5'-alkylene phosphonates, chiral phosphonates, phosphinates, phosphoramidates, such as 3'-amino phosphoramidates and aminoalkyl phosphoramidates, phosphorodiamidates, thionophosphoramidates, thionoalkyl phosphonates, thionoalkyl phosphotriesters, selenophosphates, and boranophosphates having normal 3'-5' linkages, 2'-5' linked analogs, and those with reverse polarity, where one or more internucleotide linkages are 3' to 3', 5' to 5', or 2' to 2' linkages.

[0032] Nucleic acids may contain polynucleotide backbones formed by short chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more heteroatom or heterocyclic internucleoside linkages, including those with morpholino linkages (formed in part from the sugar portion of the nucleoside), siloxane backbones, sulfide, sulfoxide, and sulfone backbones, formacetyl and thioformacetyl backbones, methyleneformacetyl and thioformacetyl backbones, riboacetyl backbones, alkene-containing backbones, sulfamate backbones, methyleneimino and methylenehydrazino backbones, sulfonate and sulfonamide backbones, amide backbones, and others with mixed N, O, S, and CH2 component moieties.

[0033] Nucleic acids may include nucleic acid mimetics. The term "mimetics" may be intended to include polynucleotides in which only the furanose ring or both the furanose ring and the internucleotide linkage are replaced with non-furanose groups, and replacement of only the furanose ring may also be referred to as a sugar surrogate. The heterocyclic base moiety or modified heterocyclic base moiety may be maintained for hybridization with an appropriate target nucleic acid. One such nucleic acid may be a peptide nucleic acid (PNA). In a PNA, the sugar backbone of a polynucleotide may be replaced with an amide-containing backbone, in particular an aminoethylglycine backbone. The nucleotides may be retained and are directly or indirectly bound to the aza nitrogen atoms of the amide portion of the backbone. The backbone in a PNA compound may contain two or more linked aminoethylglycine units, thereby providing the PNA with an amide-containing backbone. The heterocyclic base moiety may be directly or indirectly bound to the aza nitrogen atoms of the amide portion of the backbone. The nucleic acid may include a morpholino backbone structure. For example, the nucleic acid may include a six-membered morpholino ring instead of a ribose ring. In some of these embodiments, phosphorodiamidates or other non-phosphodiester internucleoside linkages may replace the phosphodiester linkages.

[0034] Nucleic acids may include linked morpholino units (e.g., morpholino nucleic acids) having heterocyclic bases attached to the morpholino ring. Linking groups may link the morpholino monomer units in the morpholino nucleic acid. Non-ionic morpholino-based oligomeric compounds may have less undesirable interactions with intracellular proteins. Morpholino-based polynucleotides may be non-ionic mimics of nucleic acids. Various compounds within the morpholino class may be linked using different linking groups. A further class of polynucleotide mimics may be referred to as cyclohexenyl nucleic acids (CeNA). The furanose ring normally present in nucleic acid molecules may be replaced with a cyclohexenyl ring. CeNA DMT-protected phosphoramidite monomers may be prepared and used in the synthesis of oligomeric compounds using phosphoramidite chemistry. Incorporation of CeNA monomers into nucleic acid strands may increase the stability of DNA / RNA hybrids. CeNA oligoadenylates can form complexes with nucleic acid complements with stability similar to the natural complexes. Further modifications can include locked nucleic acids (LNAs) in which a 2'-hydroxyl group is linked to the 4' carbon atom of the sugar ring, thereby forming a 2'-C,4'-C-oxymethylene linkage, thereby forming a bicyclic sugar moiety. The linkage can be a methylene (-CH2) group bridging the 2' oxygen atom and the 4' carbon atom, where n is 1 or 2. LNAs and LNA analogs can exhibit very high duplex thermal stability with complementary nucleic acids (Tm=+3 to +10°C), stability against 3'-exonuclease degradation, and good solubility properties.

[0035] Nucleic acids can also include nucleobase (often simply referred to as "base") modifications or substitutions. As used herein, "unmodified" or "natural" nucleobases can include purine bases (e.g., adenine (A) and guanine (G)) and pyrimidine bases (e.g., thymine (T), cytosine (C), and uracil (U)). Modified nucleobases include other synthetic and natural nucleobases, such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C=C-CH3) uracil and cytosine and other alkyl derivatives of the pyrimidine bases, 6-azouracil, cytosine ... Included are tosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl, and other 8-substituted adenines and guanines, 5-halo, particularly 5-bromo, 5-trifluoromethyl, and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-aminoadenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine, and 3-deazaguanine and 3-deazaadenine.Modified nucleobases include tricyclic pyrimidines, such as phenoxazine cytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), G-clamps, such as substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), G-clamps, such as substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), carbazole cytidines (2H-pyrimido(4,5-b)indol-2-one), pyridoindol cytidines (H-pyrido(3',2':4,5)pyrrolo[2,3-d]pyrimidin-2-one).

[0036] As used herein, the term "sample" may refer to a composition that contains a target. Samples suitable for analysis by the disclosed methods, devices, and systems include cells, tissues, organs, or organisms. As used herein, the term "sample collection device" or "device" may refer to a device capable of collecting a section of a sample and / or placing the section on a substrate. A sample device may refer to, for example, a fluorescence activated cell sorter (FACS) machine, a cell sorter, a biopsy needle, a biopsy device, a tissue sectioning device, a microfluidic device, a blade grid, and / or a microtome. As used herein, the term "solid support" may refer to a discrete solid or semi-solid surface to which a plurality of barcodes (e.g., stochastic barcodes) may be attached. A solid support may encompass any type of solid, porous, or hollow sphere, ball, bearing, cylinder, or other similar configuration composed of plastic, ceramic, metal, or polymeric material (e.g., hydrogel) to which a nucleic acid may be immobilized (e.g., covalently or non-covalently). A solid support may include discrete particles that may be spherical (e.g., microspheres) or may have a non-spherical or irregular shape, such as a cube, cube-like, pyramidal, cylindrical, conical, rectangular, or discoid shape. A bead may be of a non-spherical shape. A plurality of solid supports spaced apart in an array may not include a substrate. A solid support may be used interchangeably with the term "beads."

[0037] As used herein, the term "stochastic barcode" may refer to a polynucleotide sequence that includes a label of the present disclosure. A stochastic barcode may be a polynucleotide sequence that can be used for stochastic barcoding. A stochastic barcode may be used to quantify a target in a sample. A stochastic barcode may be used to control errors that may occur after associating a label with a target. For example, a stochastic barcode may be used to evaluate amplification or sequencing errors. A stochastic barcode associated with a target may be referred to as a stochastic barcode-target or a stochastic barcode-tag-target. As used herein, the term "gene-specific stochastic barcode" may refer to a polynucleotide sequence that includes a label and a target binding region that is gene-specific. A stochastic barcode may be a polynucleotide sequence that can be used for stochastic barcoding. A stochastic barcode may be used to quantify a target in a sample. A stochastic barcode may be used to control errors that may occur after associating a label with a target. For example, a stochastic barcode may be used to evaluate amplification or sequencing errors. A stochastic barcode associated with a target may be referred to as a stochastic barcode-target or a stochastic barcode-tag-target.

[0038] As used herein, the term "stochastic barcoding" may refer to random labeling (e.g., barcoding) of nucleic acids. Stochastic barcoding can utilize a Poisson recursive strategy to associate labels and quantify the labels associated with targets. As used herein, the term "stochastic barcoding" can be used interchangeably with "stochastic labeling." As used herein, the term "target" may refer to a composition that may be associated with a barcode (e.g., a stochastic barcode). Exemplary targets suitable for analysis by the methods, devices, and systems of the present disclosure include oligonucleotides, DNA, RNA, mRNA, microRNA, tRNA, and the like. Targets may be single-stranded or double-stranded. In some embodiments, targets may be proteins, peptides, or polypeptides. In some embodiments, targets are lipids. As used herein, "target" may be used interchangeably with "species."

[0039] As used herein, the term "reverse transcriptase" may refer to a group of enzymes that have reverse transcriptase activity (i.e., catalyze the synthesis of DNA from an RNA template). In general, such enzymes include, but are not limited to, retroviral reverse transcriptases, retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retron reverse transcriptases, bacterial reverse transcriptases, group II intron-derived reverse transcriptases, and mutants, variants, or derivatives thereof. Non-retroviral reverse transcriptases include non-LTR retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retron reverse transcriptases, and group II intron reverse transcriptases. Examples of group II intron reverse transcriptases include Lactococcus lactis LI.LtrB intron reverse transcriptase, Thermosynechococcus elongatus TeI4c intron reverse transcriptase, or Geobacillus stearothermophilus GsI-IIC intron reverse transcriptase. Other classes of reverse transcriptases can include the numerous classes of non-retroviral reverse transcriptases (i.e., retrons, group II introns, and diversity generating retroelements, among others).

[0040] The terms "universal adapter primer", "universal primer adapter" or "universal adapter sequence" are used interchangeably to refer to a nucleotide sequence that can hybridize to a barcode (e.g., a stochastic barcode) and can be used to generate a gene-specific barcode. The universal adapter sequence can be, for example, a known sequence that is universal across all barcodes used in the methods of the present disclosure. For example, when multiple targets are labeled using the methods disclosed herein, each of the target-specific sequences can be linked to the same universal adapter sequence. In some embodiments, more than one universal adapter sequence can be used in the methods disclosed herein. For example, when multiple targets are labeled using the methods disclosed herein, at least two of the target-specific sequences are linked to different universal adapter sequences. The universal adapter primer and its complement can be included in two oligonucleotides, one of which includes the target-specific sequence and the other of which includes the barcode. For example, the universal adapter sequence can be part of an oligonucleotide that includes a target-specific sequence to generate a nucleotide sequence that is complementary to the target nucleic acid. A second oligonucleotide comprising the barcode and a complementary sequence of the universal adapter sequence may hybridize to the nucleotide sequence to generate a target-specific barcode (target-specific stochastic barcode). In some embodiments, the universal adapter primer has a different sequence than the universal PCR primer used in the disclosed methods.

[0041] Barcode Barcoding, e.g., probabilistic barcoding, is described, for example, in U.S. Patent Application Publication No. US2015 / 0299784, International Publication No. WO2015 / 031691, and Fu et al, Proc Natl Acad Sci USA 2011 May 31;108(22):9026-31, the contents of which are incorporated herein in their entirety. In some embodiments, the barcodes disclosed herein can be probabilistic barcodes, which can be polynucleotide sequences that can be used to stochastically label (e.g., barcode, tag) targets. A barcode may be referred to as a stochastic barcode if the ratio of the number of distinct barcode sequences of the stochastic barcode to the number of occurrences of any of the targets to be labeled can be 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or a number or range between any two of these values, or can be approximately these values ​​or such number or range. The targets may be mRNA species that include mRNA molecules with identical or nearly identical sequences. A barcode may be referred to as a stochastic barcode if the ratio of the number of distinct barcode sequences of the stochastic barcode to the number of occurrences of any of the targets to be labeled is at least or at most 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1. The barcode sequences of the stochastic barcode may be referred to as molecular labels.

[0042] A barcode, e.g., a stochastic barcode, may include one or more labels. Exemplary labels may include a universal label, a cell label, a barcode sequence (e.g., a molecular label), a sample label, a plate label, a spatial label, and / or a pre-spatial label. FIG. 1 shows an exemplary barcode 104 having a spatial label. The barcode 104 may include a 5' amine that allows the barcode to be linked to a solid support 105. The barcode may include a universal label, a dimension label, a spatial label, a cell label, and / or a molecular label. The order of the different labels (including but not limited to the universal label, the dimensional label, the spatial label, the cell label, and the molecular label) in the barcode may vary. For example, as shown in FIG. 1, the universal label may be the 5'-most label and the molecular label may be the 3'-most label. The spatial label, the dimensional label, and the cell label may be in any order. In some embodiments, the universal label, the spatial label, the dimensional label, the cell label, and the molecular label are in any order. The barcode may include a target binding region. The target binding region may interact with a target (e.g., a target nucleic acid, RNA, mRNA, DNA) in a sample. For example, the target binding region may include an oligo(dT) sequence that can interact with the poly(A) tail of an mRNA. In some cases, the labels of the barcode (e.g., the universal label, the dimensional label, the spatial label, the cellular label, and the barcode sequence) may be spaced 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides apart.

[0043] Labels, e.g., cell labels, may include a set of unique nucleic acid subsequences of defined length, e.g., 7 nucleotides each (equivalent to the number of bits used in some Hamming error correction codes), that may be designed to provide error correction capabilities. An error correction subsequence set that includes 7 nucleotide sequences may be designed such that any pairwise combination of sequences in the set exhibits a defined "genetic distance" (or number of mismatched bases), e.g., an error correction subsequence set may be designed to exhibit a genetic distance of 3 nucleotides. In this case, consideration of the error correction sequences in the sequence data set of the labeled target nucleic acid molecule (described in more detail below) may allow for the detection or correction of amplification or sequencing errors. In some embodiments, the length of the nucleic acid subsequences used to create the error correction code may vary, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 31, 40, 50 nucleotides, or a number or range between any two of these values, or approximately these values ​​or such number or range of nucleotides in length. In some embodiments, nucleic acid subsequences of other lengths may be used to create the error correcting code.

[0044] The barcode may include a target binding region. The target binding region may interact with a target in the sample. The target may be or include ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNA each including a poly(A) tail, or any combination thereof. In some embodiments, the multiple targets may include deoxyribonucleic acid (DNA).

[0045] In some embodiments, the target binding region may include an oligo(dT) sequence that can interact with the poly(A) tail of mRNA. One or more of the labels of the barcode (e.g., universal label, dimensional label, spatial label, cellular label, and barcode sequence (e.g., molecular label)) may be spaced from another one or two of the remaining labels of the barcode by a spacer. The spacer may be, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides, or more. In some embodiments, none of the labels of the barcode are spaced apart by a spacer.

[0046] Universal Signage The barcode may include one or more universal labels. In some embodiments, the one or more universal labels may be the same for all barcodes in a set of barcodes bound to a given solid support. In some embodiments, the one or more universal labels may be the same for all barcodes bound to a plurality of beads. In some embodiments, the universal label may include a nucleic acid sequence that can hybridize to a sequencing primer. The sequencing primer can be used to sequence the barcode that includes the universal label. The sequencing primer (e.g., a universal sequencing primer) may include a sequencing primer associated with a high-throughput sequencing platform. In some embodiments, the universal label may include a nucleic acid sequence that can hybridize to a PCR primer. In some embodiments, the universal label may include a nucleic acid sequence that can hybridize to a sequencing primer and a PCR primer. The nucleic acid sequence of the universal label that can hybridize to a sequencing primer or a PCR primer may be referred to as a primer binding site. The universal label may include a sequence that can be used to initiate transcription of the barcode. The universal label may include a sequence that can be used to extend the barcode or a region within the barcode. The universal label may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range between any two of these values, or about these values ​​or such number or range of nucleotides. For example, the universal label may include at least about 10 nucleotides. The universal label may be, for example, at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length. In some embodiments, a cleavable linker or modified nucleotide may be part of the universal label sequence to allow the barcode to be cleaved from the support.

[0047] dimensional indicator A barcode may include one or more dimensional labels. In some embodiments, a dimensional label may include a nucleic acid sequence that provides information about the dimension in which labeling (e.g., stochastic labeling) occurred. For example, a dimensional label may provide information about the time a target was barcoded. A dimensional label may be associated with the time of barcoding (e.g., stochastic barcoding) in a sample. A dimensional label may be activated at the time of labeling. Different dimensional labels may be activated at different times. A dimensional label provides information about the order in which targets, groups of targets, and / or samples were barcoded. For example, a cell population may be barcoded in the G0 phase of the cell cycle. Cells may be pulsed again with a barcode (e.g., a stochastic barcode) in the G1 phase of the cell cycle. Cells may be pulsed again with a barcode in the S phase of the cell cycle, and so on. The barcodes in each pulse (e.g., each stage of the cell cycle) may include different dimensional labels. In this way, the dimensional labels provide information about which targets were labeled at which stage of the cell cycle. Dimensional labeling can examine many different biological times. Exemplary biological times can include, but are not limited to, cell cycle, transcription (e.g., transcription initiation), and transcript degradation. In another example, a sample (e.g., a cell, a cell population) can be stochastically labeled before and / or after treatment with a drug and / or therapy. Changes in copy number of distinct targets can indicate the response of the sample to the drug and / or therapy.

[0048] The dimension label may be activatable. The activatable dimension label may be activated at a particular time. The activatable label may, for example, be continuously activated (e.g., not turned off). The activatable dimension label may, for example, be reversibly activatable (e.g., the activatable dimension label may be turned on and off). The dimension label may, for example, be reversibly activatable at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more times. The dimension label may, for example, be reversibly activatable at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more times. In some embodiments, the dimension label may be activated by fluorescence, light, chemical events (e.g., cleavage, ligation of another molecule, addition of a modification (e.g., pegylation, sumoylation, acetylation, deacetylation, demethylation), photochemical events (e.g., photocaging), and introduction of unnatural nucleotides.

[0049] The dimension label, in some embodiments, may be the same for all barcodes (e.g., stochastic barcodes) bound to a given solid support (e.g., a bead), but may be different for different solid supports (e.g., beads). In some embodiments, at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100% of the barcodes on the same solid support may include the same dimension label. In some embodiments, at least 60% of the barcodes on the same solid support may include the same dimension label. In some embodiments, at least 95% of the barcodes on the same solid support may include the same dimension label. On multiple solid supports (e.g., beads), 6Many unique dimension label sequences, even more than 10, may be presented. A dimension label may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range between any two of these values, or about these values ​​or such number or range of nucleotides. A dimension label may be at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length. A dimension label may comprise from about 5 to about 200 nucleotides. A dimension label may comprise from about 10 to about 150 nucleotides. A dimension label may comprise from about 20 to about 125 nucleotides in length.

[0050] spatial sign The barcode may include one or more spatial labels. In some embodiments, the spatial label may include a nucleic acid sequence that provides information about the spatial orientation of the target molecule associated with the barcode. The spatial label may be associated with a coordinate in the sample. The coordinate may be a fixed coordinate. For example, the coordinate may be fixed relative to a substrate. The spatial label may be referenced to a two-dimensional or three-dimensional grid. The coordinate may be fixed relative to a landmark. The landmark may be identifiable in space. The landmark may be a structure that can be imaged. The landmark may be a biological structure, e.g., an anatomical landmark. The landmark may be a cellular landmark, e.g., an organelle. The landmark may be a non-natural landmark, e.g., a structure with an identifiable identifier, e.g., a color code, a bar code, a magnetic property, a fluorescent property, a radioactive property, or a unique size or shape. The spatial label may be associated with a physical compartment (e.g., a well, a container, or a droplet). In some embodiments, multiple spatial labels are used together to code one or more locations in space.

[0051] The spatial labels may be the same for all barcodes bound to a given solid support (e.g., beads), but may be different for different solid supports (e.g., beads). In some embodiments, the percentage of barcodes containing the same spatial label on the same solid support may be 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values, or may be approximately these values ​​or such numbers or ranges. In some embodiments, the percentage of barcodes containing the same spatial label on the same solid support may be at least or at most 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. In some embodiments, at least 60% of the barcodes on the same solid support may contain the same spatial label. In some embodiments, at least 95% of the barcodes on the same solid support may contain the same spatial label.

[0052] On multiple solid supports (e.g., beads), 6 Many unique spatial marker sequences, even more than 10, may be presented. Spatial markers can be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range between any two of these values, or about these values ​​or such number or range of nucleotides. Spatial markers can be, for example, at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length. Spatial markers can include between about 5 and about 200 nucleotides. Spatial markers can include between about 10 and about 150 nucleotides. Spatial markers can include between about 20 and about 125 nucleotides in length.

[0053] cell labeling A barcode (e.g., a stochastic barcode) may include one or more cell labels. In some embodiments, the cell labels may include a nucleic acid sequence that provides information to determine which target nucleic acid originated from which cell. In some embodiments, the cell labels are the same for all barcodes bound to a given solid support (e.g., a bead), but different for different solid supports (e.g., beads). In some embodiments, the percentage of barcodes containing the same cell label on the same solid support may be 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values, or approximately these values ​​or such number or range. In some embodiments, the percentage of barcodes containing the same cell label on the same solid support may be 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%, or approximately these values. For example, at least 60% of the barcodes on the same solid support may contain the same cell label. As another example, at least 95% of the barcodes on the same solid support may contain the same cell label. On multiple solid supports (e.g., beads), 6 Many unique cell marker sequences, even more than 10, may be presented. The cell marker can be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range between any two of these values, or about these values ​​or such number or range of nucleotides. The cell marker can be, for example, at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length. For example, the cell marker can include between about 5 and about 200 nucleotides. As another example, the cell marker can include about 10 to about 150 nucleotides. As yet another example, the cell marker can include about 20 to about 125 nucleotides in length.

[0054] Barcode sequence The barcode may include one or more barcode sequences. In some embodiments, the barcode sequence may include a nucleic acid sequence that provides information about the specific type of target nucleic acid species that is hybridized to the barcode. The barcode sequence includes a nucleic acid sequence that provides a counter (e.g., provides a rough approximation) for the specific occurrence of the target nucleic acid species that is hybridized to the barcode (e.g., the target binding region).

[0055] In some embodiments, a diverse set of barcode sequences is attached to a given solid support (e.g., a bead). 2 pieces, 10 3 pieces, 10 4 pieces, 10 5 pieces, 10 6 pieces, 10 7 pieces, 10 8 pieces, 10 9 There may be at least about 10 unique molecular label sequences, or a number or range between or about any two of these values. For example, the plurality of barcodes may include about 6561 barcode sequences having distinct sequences. As another example, the plurality of barcodes may include about 65536 barcode sequences having distinct sequences. In some embodiments, there may be at least about 10 unique molecular label sequences, or a number or range between or about any two of these values. 2 pieces, 10 3 pieces, 10 4 pieces, 10 5 pieces, 10 6 pieces, 10 7 pieces, 10 8 Pieces or 10 9 There may be one unique barcode sequence. The unique molecular label sequence may be attached to a given solid support (e.g., a bead). In some embodiments, the unique molecular label sequence is partially or entirely encompassed by a particle (e.g., a hydrogel bead).

[0056] The length of the barcode may vary in different implementations. For example, the barcode may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or a number or range between any two of these values, or may be about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or a number or range between any two of these values, in nucleotide length. As another example, the barcode may be at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides long.

[0057] molecular label A barcode (e.g., a stochastic barcode) may include one or more molecular labels. A molecular label may include a barcode sequence. In some embodiments, a molecular label may include a nucleic acid sequence that provides identification of information regarding a particular type of target nucleic acid species hybridized to the barcode. A molecular label includes a nucleic acid sequence that provides a counter for the particular occurrence of a target nucleic acid species hybridized to the barcode (e.g., a target binding region). In some embodiments, a diverse set of molecular labels is attached to a given solid support (e.g., a bead). 2 pieces, 10 3 pieces, 10 4 pieces, 10 5 pieces, 10 6 pieces, 10 7 pieces, 10 8 pieces, 10 9 or a number or range between any two of these values, or about 10 2 pieces, 10 3 pieces, 10 4 pieces, 10 5 pieces, 10 6 pieces, 10 7 pieces, 10 8 pieces, 10 9There may be at least about 10 unique molecular label sequences, or a number or range between any two of these values. For example, the plurality of barcodes may include about 6561 molecular labels with distinct sequences. As another example, the plurality of barcodes may include about 65536 molecular labels with distinct sequences. In some embodiments, there may be at least about 10 unique molecular label sequences, or a number or range between any two of these values. 2 pieces, 10 3 pieces, 10 4 pieces, 10 5 pieces, 10 6 pieces, 10 7 pieces, 10 8 Pieces or 10 9 There may be one unique molecular label sequence. A barcode with a unique molecular label sequence may be attached to a given solid support (e.g., a bead).

[0058] For barcoding using multiple stochastic barcodes (e.g., stochastic barcoding), the ratio of the number of distinct molecular label sequences to the number of occurrences of any of the targets can be 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or a number or range between any two of these values, or may be about 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or a number or range between any two of these values. The target may be an mRNA species that includes mRNA molecules with identical or nearly identical sequences. In some embodiments, the ratio of the number of different molecular label sequences to the number of occurrences of any of the targets is at least or at most 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100:1.

[0059] Molecular labels can be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range between any two of these values, or a number or range of nucleotides approximately equal to or equal to these values ​​or such number or range. Molecular labels can be, for example, at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length.

[0060] Target binding region The barcode may include one or more target binding regions, e.g., capture probes. In some embodiments, the target binding region may hybridize to a target of interest. In some embodiments, the target binding region may include a nucleic acid sequence that specifically hybridizes to a target (e.g., a target nucleic acid, a target molecule, e.g., a cellular nucleic acid to be analyzed), e.g., a specific gene sequence. In some embodiments, the target binding region may include a nucleic acid sequence that can bind (e.g., hybridize) to a specific position of a specific target nucleic acid. In some embodiments, the target binding region may include a nucleic acid sequence that is capable of specific hybridization to a restriction enzyme site overhang (e.g., an EcoRI sticky end overhang). The barcode can then be ligated to any nucleic acid molecule that includes a sequence complementary to the restriction site overhang.

[0061] In some embodiments, the target binding region may include a non-specific target nucleic acid sequence. A non-specific target nucleic acid sequence may refer to a sequence that can bind to multiple target nucleic acids independent of the specific sequence of the target nucleic acid. For example, the target binding region may include a random multimer sequence, or an oligo(dT) sequence that hybridizes to a poly(A) tail on an mRNA molecule. The random multimer sequence may be, for example, a random dimer, trimer, tetramer, pentamer, hexamer, heptamer, octamer, nonamer, decamer, or any longer multimer sequence of any length. In some embodiments, the target binding region is the same for all barcodes bound to a given bead. In some embodiments, the target binding regions of multiple barcodes bound to a given bead may include two or more different target binding sequences. A target binding region can be 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range between any two of these values, or about these values ​​or such number or range of nucleotides in length. A target binding region can be up to about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or more.

[0062] In some embodiments, the target binding region may comprise an oligo(dT) that can hybridize with an mRNA that comprises a polyadenylated end. The target binding region may be gene-specific. For example, the target binding region may be configured to hybridize to a specific region of the target. The target binding region may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 27, 28, 29, 30, or a number or range between any two of these values, or approximately these values ​​or such number or range of nucleotides in length. The target binding region may be at least or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 27, 28, 29, or 30 nucleotides in length. The target binding region may be about 5-30 nucleotides in length. When the barcode comprises a gene-specific target binding region, the barcode may be referred to herein as a gene-specific barcode.

[0063] Orientation Characteristics A stochastic barcode (e.g., a stochastic barcode) may include one or more orientation properties that can be used to orient (e.g., align) the barcode. The barcode may include moieties for isoelectric focusing. Different barcodes may include different isoelectric focusing points. When these barcodes are introduced into a sample, the sample may undergo isoelectric focusing to orient the barcodes in a known manner. In this manner, the orientation properties may be used to develop a known map of the barcodes in the sample. Exemplary orientation properties may include electrophoretic mobility (e.g., based on the size of the barcode), isoelectric point, spin, conductivity, and / or self-assembly. For example, a barcode with an orientation property of self-assembly may self-assemble into a particular orientation (e.g., a nucleic acid nanostructure) when activated.

[0064] affinity properties A barcode (e.g., a stochastic barcode) may include one or more affinity features. For example, a spatial label may include an affinity feature. An affinity feature may include a chemical and / or biological moiety that can facilitate binding of the barcode to another entity (e.g., a cellular receptor). For example, an affinity feature may include an antibody, e.g., an antibody specific to a particular moiety (e.g., a receptor) on a sample. In some embodiments, an antibody may direct a barcode to a particular cell type or molecule. Targets at and / or near a particular cell type or molecule may be labeled (e.g., stochastically labeled). An affinity feature may provide spatial information in addition to the nucleotide sequence of the spatial label, since in some embodiments, an antibody may direct a barcode to a particular location. An antibody may be a therapeutic antibody, e.g., a monoclonal or polyclonal antibody. An antibody may be humanized or chimeric. An antibody may be a naked antibody or a fusion antibody.

[0065] An antibody can be a full-length (i.e., naturally occurring or generated by conventional immunoglobulin gene fragment recombination processes) immunoglobulin molecule (e.g., an IgG antibody), or an immunologically active (i.e., specific binding) portion of an immunoglobulin molecule, such as an antibody fragment. An antibody fragment can be, for example, a portion of an antibody, such as F(ab')2, Fab', Fab, Fv, sFv, etc. In some embodiments, an antibody fragment can bind to the same antigen recognized by a full-length antibody. Antibody fragments can include isolated fragments of the variable regions of an antibody, such as "Fv" fragments consisting of the variable regions of the heavy and light chains, and recombinant single-chain polypeptide molecules in which the variable regions of the light and heavy chains are connected by a peptide linker ("scFv protein"). Exemplary antibodies can include, but are not limited to, antibodies against cancer cells, antibodies against viruses, antibodies that bind to cell surface receptors (CD8, CD34, CD45), and therapeutic antibodies.

[0066] Universal Adapter Primer A barcode may include one or more universal adapter primers. For example, a gene-specific barcode, such as a gene-specific stochastic barcode, may include a universal adapter primer. A universal adapter primer may refer to a nucleotide sequence that is universal across all barcodes. A universal adapter primer can be used to construct a gene-specific barcode. A universal adapter primer may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 27, 28, 29, 30 nucleotides in length, or a number or range between any two of these, or approximately these values ​​or such number or range of nucleotides. The universal adapter primer can be at least or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 27, 28, 29, or 30 nucleotides in length. The universal adapter primer can be between 5 and 30 nucleotides in length.

[0067] Linker When a barcode includes more than one type of label (e.g., more than one cell label or more than one barcode sequence, e.g., one molecular label), the labels may be interspersed with linker label sequences. The linker label sequence may be at least about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or more nucleotides in length. The linker label sequence may be up to about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or more nucleotides in length. In some cases, the linker label sequence is 12 nucleotides in length. The linker label sequence may be used to facilitate synthesis of the barcode. The linker label may include an error-correcting (e.g., Hamming) code.

[0068] solid support The barcodes disclosed herein, e.g., stochastic barcodes, may be associated with a solid support in some embodiments. The solid support may be, for example, a synthetic particle. In some embodiments, some or all of the barcode sequences, e.g., molecular labels of stochastic barcodes (e.g., a first barcode sequence) of a plurality of barcodes (e.g., a first plurality of barcodes) on a solid support, differ by at least one nucleotide. The cellular labels of barcodes on the same solid support may be the same. The cellular labels of barcodes on different solid supports may differ by at least one nucleotide. For example, a first cellular label of a first plurality of barcodes on a first solid support may have the same sequence and a second cellular label of a second plurality of barcodes on a second solid support may have the same sequence. The first cellular label of the first plurality of barcodes on a first solid support and the second cellular label of the second plurality of barcodes on a second solid support may differ by at least one nucleotide. The cellular labels may be, for example, about 5-20 nucleotides in length. The barcode sequence can be, for example, about 5-20 nucleotides in length. The synthetic particle can be, for example, a bead.

[0069] The beads can be, for example, silica gel beads, controlled pore glass beads, magnetic beads, Dynabeads, sephadex / sepharose beads, cellulose beads, polystyrene beads, or any combination thereof. The beads can include materials such as polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogels, paramagnetic materials, ceramics, plastics, glass, methylstyrene, acrylic polymers, titanium, latex, sepharose, cellulose, nylon, silicone, or any combination thereof. In some embodiments, the beads can be polymer beads, such as deformable beads or gel beads (e.g., gel beads from 10X Genomics, San Francisco, CA) functionalized with barcodes or stochastic barcodes. In some implementations, the gel beads can include polymer-based gels. The gel beads can be generated, for example, by encapsulating one or more polymer precursors in a droplet. The gel beads can be generated when the polymer precursors are exposed to an accelerant (e.g., tetramethylethylenediamine (TEMED)).

[0070] In some embodiments, the particles may be disintegrable (e.g., dissolvable, degradable). For example, polymer beads may, for example, dissolve, melt, or degrade under desired conditions. The desired conditions may include environmental conditions. The desired conditions may result in the dissolution, melt, or decomposition of the polymer beads in a controlled manner. Gel beads may dissolve, melt, or decompose due to chemical, physical, biological, thermal, magnetic, electrical, optical, or any combination thereof.

[0071] Analytes and / or reagents, e.g., oligonucleotide barcodes, may be linked / immobilized, for example, to the interior surface of the gel bead (e.g., the interior accessible through diffusion of the oligonucleotide barcodes and / or the material used to generate the oligonucleotide barcodes) and / or to the exterior surface of the gel bead or any other microcapsule described herein. Linkage / immobilization may be via any form of chemical bond (e.g., covalent bond, ionic bond) or physical phenomenon (e.g., van der Waals forces, dipole-dipole interactions, etc.). In some embodiments, linkage / immobilization of reagents to gel beads or any other microcapsule described herein may be reversible, such as, for example, via a labile moiety (e.g., via a chemical crosslinker, including chemical crosslinkers described herein). Upon application of a stimulus, the labile moiety may be cleaved and the immobilized reagent released. In some embodiments, the labile moiety is a disulfide bond. For example, in cases where oligonucleotide barcodes are immobilized to gel beads via disulfide bonds, exposing the disulfide bond to a reducing agent may cleave the disulfide bond and release the oligonucleotide barcode from the bead. The labile moiety may be included as part of the gel bead or microcapsule, as part of a chemical linker that links the reagent or analyte to the gel bead or microcapsule, and / or as part of the reagent or analyte. In some embodiments, at least one barcode of the plurality of barcodes may be immobilized to the particle, partially immobilized to the particle, encapsulated in the particle, partially encapsulated in the particle, or any combination thereof.

[0072] In some embodiments, the gel beads may comprise a wide variety of different polymers, including, but not limited to, polymers, thermosensitive polymers, light sensitive polymers, magnetic polymers, pH sensitive polymers, salt sensitive polymers, chemically sensitive polymers, polyelectrolytes, polysaccharides, peptides, proteins, and / or plastics. Polymers can include, but are not limited to, materials such as poly(N-isopropylacrylamide) (PNIPAAm), poly(styrenesulfonate) (PSS), poly(allylamine) (PAAm), poly(acrylic acid) (PAA), poly(ethyleneimine) (PEI), poly(diallyldimethyl-ammonium chloride) (PDADMAC), poly(pyrrole) (PPy), poly(vinylpyrrolidone) (PVPON), poly(vinylpyridine) (PVP), poly(methacrylic acid) (PMAA), poly(methyl methacrylate) (PMMA), polystyrene (PS), poly(tetrahydrofuran) (PTHF), poly(phthalaldehyde) (PPA), poly(hexylviologen) (PHV), poly(L-lysine) (PLL), poly(L-arginine) (PARG), and poly(lactic-co-glycolic acid) (PLGA).

[0073] A number of chemical stimuli can be used to trigger the collapse, dissolution, or disintegration of beads. Examples of these chemical changes can include, but are not limited to, pH-mediated changes to the bead wall, collapse of the bead wall via chemical cleavage of cross-links, triggering depolymerization of the bead wall, and switching reactions of the bead wall. Bulk changes can also be used to trigger the collapse of the beads. Bulk or physical changes to microcapsules through various stimuli also offer many advantages in designing capsules to release reagents. Bulk or physical changes occur on a macroscopic scale, with bead bursting being the result of mechanical-physical forces induced by the stimuli. These processes can include, but are not limited to, pressure-induced bursting, melting of the bead wall, or changes in the porosity of the bead wall.

[0074] Biological stimuli can also be used to trigger the collapse, dissolution, or degradation of beads. In general, biological triggers resemble chemical triggers, but many examples use biomolecules, or molecules commonly found in biological systems, such as enzymes, peptides, sugars, fatty acids, nucleic acids, and the like. For example, beads can include polymers with peptide crosslinks that are susceptible to cleavage by specific proteases. More specifically, one example can include microcapsules that include GFLGK peptide crosslinks. Upon addition of a biological trigger, such as the protease cathepsin B, the peptide crosslinks in the shell wall are cleaved, releasing the contents of the bead. In other cases, the protease can be heat activated. In another example, beads include a shell wall that includes cellulose. The addition of chitosan, a hydrolytic enzyme, serves as a biological trigger for cleavage of the cellulose bonds, depolymerization of the shell wall, and release of its internal contents.

[0075] Beads can also be induced to release their contents upon application of a thermal stimulus. A change in temperature can cause various changes in the beads. A change in heat can cause the beads to melt, such that the bead walls collapse. In some embodiments, heat can increase the internal pressure of the internal components of the beads, such that the beads collapse or explode. In some embodiments, heat can transform the beads into a compressed, dehydrated state. Heat can also act on a heat-sensitive polymer within the bead's wall, causing the beads to collapse. The inclusion of magnetic nanoparticles in the bead walls of microcapsules can trigger the collapse of beads as well as guide the beads in arrays. The devices of the present disclosure can include magnetic beads for any purpose. In one example, the incorporation of Fe3O4 nanoparticles into polyelectrolyte-containing beads triggers collapse in the presence of an oscillating magnetic field stimulus.

[0076] Beads can also disintegrate, dissolve, or break down as a result of electrical stimulation. Similar to the magnetic particles described in the previous section, electrically sensitive beads can trigger both disintegration of the beads and other functions such as alignment in an electric field, electrical conduction, or redox reactions. In one example, beads containing electrically sensitive materials are aligned in an electric field so that the release of internal reagents can be controlled. In another example, the electric field can induce redox reactions within the bead wall itself, which can increase porosity. Light stimulation can also be used to disrupt the beads. Numerous light triggers are possible, including systems that use various molecules such as nanoparticles and chromophores that can absorb photons of a specific range of wavelengths. For example, metal oxide coatings can be used as capsule triggers. UV irradiation of SiO2-coated polyelectrolyte capsules can result in the collapse of the bead walls. In yet another example, photoswitchable materials, such as azobenzene groups, can be incorporated into the bead walls. Upon application of UV or visible light, chemicals such as these absorb photons and undergo reversible cis to trans isomerization. In this embodiment, the incorporation of a light switch results in a bead wall that can collapse or become more porous upon application of a light trigger.

[0077] For example, in a non-limiting example of barcoding (e.g., stochastic barcoding) shown in FIG. 2, after cells, e.g., single cells, are introduced into a plurality of microwells of a microwell array in block 208, beads can be introduced into a plurality of microwells of the microwell array in block 212. Each microwell may contain one bead. The beads may contain multiple barcodes. The barcodes may include 5' amine regions attached to the beads. The barcodes may include a universal label, a barcode sequence (e.g., a molecular label), a target binding region, or any combination thereof.

[0078] The barcodes disclosed herein may be associated with (e.g., bound to) a solid support (e.g., a bead). The barcodes associated with the solid support may include a barcode sequence selected from a group including at least 100 or 1000 barcode sequences, each having a unique sequence. In some embodiments, different barcodes associated with a solid support may include barcodes with different sequences. In some embodiments, a percentage of the barcodes associated with a solid support include the same cell label. For example, the percentage may be 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values, or may be approximately these values ​​or such number or range. As another example, the percentage may be at least or at most 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. In some embodiments, the barcodes associated with the solid supports may have the same cell marker. The barcodes associated with different solid supports may have different cell markers selected from a group including at least 100 or 1000 cell markers having unique sequences.

[0079] The barcodes disclosed herein may be associated with (e.g., bound to) a solid support (e.g., a bead). In some embodiments, barcoding the multiple labels in the sample may be performed using a solid support comprising multiple synthetic particles associated with multiple barcodes. In some embodiments, the solid support may comprise multiple synthetic particles associated with multiple barcodes. The spatial labeling of the multiple barcodes on different solid supports may differ by at least one nucleotide. The solid support may comprise multiple barcodes, for example, in two or three dimensions. The synthetic particles may be beads. The beads may be silica gel beads, controlled pore glass beads, magnetic beads, Dynabeads, Sephadex / sepharose beads, cellulose beads, polystyrene beads, or any combination thereof. The solid support may include a polymer, matrix, hydrogel, needle array device, antibody, or any combination thereof. In some embodiments, the solid support may be free floating. In some embodiments, the solid support may be embedded in a semi-solid or solid array. The barcodes may not be associated with a solid support. The barcodes may be individual nucleotides. The barcode may be associated with the substrate.

[0080] As used herein, the terms "tethered," "attached," and "immobilized" are used interchangeably and can refer to covalent or non-covalent means for attaching a barcode to a solid support. Any of a variety of different solid supports can be used to attach pre-synthesized barcodes or as a solid support for in situ solid phase synthesis of barcodes. In some embodiments, the solid support is a bead. The bead may include one or more types of solid, porous, or hollow spheres, balls, bearings, cylinders, or other similar configurations on which nucleic acids can be immobilized (e.g., covalently or non-covalently). The bead may be composed of, for example, plastic, ceramic, metal, polymeric materials, or any combination thereof. The bead may be or include discrete particles that are spherical (e.g., microspheres), or may have a non-spherical or irregular shape, such as, for example, a cube, cube-like, pyramidal, cylindrical, conical, rectangular, or discoid. In some embodiments, the bead may be non-spherical in shape.

[0081] The beads may comprise a variety of materials, including, but not limited to, paramagnetic materials (e.g., magnesium, molybdenum, lithium, and tantalum), superparamagnetic materials (e.g., ferrite (Fe3O4, magnetite) nanoparticles), ferromagnetic materials (e.g., iron, nickel, cobalt, some alloys thereof, and some rare earth metal compounds), ceramic, plastic, glass, polystyrene, silica, methylstyrene, acrylic polymers, titanium, latex, sepharose, agarose, hydrogels, polymers, cellulose, nylon, or any combination thereof. In some embodiments, the beads (e.g., the beads to which the labels are attached) are hydrogel beads. In some embodiments, the beads comprise a hydrogel.

[0082] Some embodiments disclosed herein include one or more particles (e.g., beads). Each of the particles may include a plurality of oligonucleotides (e.g., barcodes). Each of the plurality of oligonucleotides may include a barcode sequence (e.g., a molecular label sequence), a cell label, and a target binding region (e.g., an oligo(dT) sequence, a gene-specific sequence, a random multimer, or a combination thereof). The cell label sequence of each of the plurality of oligonucleotides may be the same. The cell label sequences of the oligonucleotides on different particles may be different so that the oligonucleotides on different particles can be identified. The number of different cell label sequences may be different in different implementations. In some embodiments, the number of cell labeling sequences is 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 pieces, 10 7 pieces, 10 8 pieces, 10 9, a number or range between any two of these values, or more, may be about, at least, or up to these values ​​or such number or range. In some embodiments, no more than 1, 2 or less, 3 or less, 4 or less, 5 or less, 6 or less, 7 or less, 8 or less, 9 or less, 10 or less, 20 or less, 30 or less, 40 or less, 50 or less, 60 or less, 70 or less, 80 or less, 90 or less, 100 or less, 200 or less, 300 or less, 400 or less, 500 or less, 600 or less, 700 or less, 800 or less, 900 or less, 1000 or less, or more of the plurality of particles comprise oligonucleotides having the same cellular sequence. In some embodiments, the plurality of particles comprising oligonucleotides with the same cellular sequence may be up to 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10% or more. In some embodiments, none of the plurality of particles have the same cellular labeling sequence.

[0083] The oligonucleotides on each particle may include different barcode sequences (e.g., molecular labels). In some embodiments, the number of barcode sequences may be 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 pieces, 10 7 pieces, 10 8 pieces, 10 9The number of oligonucleotides may be about, at least, or at most two of these values. For example, at least 100 of the plurality of oligonucleotides include different barcode sequences. As another example, in a single particle, at least 100, 500, 1000, 5000, 10000, 15000, 20000, 50000, a number or range between any two of these values, or more of the plurality of oligonucleotides include different barcode sequences. Some embodiments provide a plurality of particles that include barcodes. In some embodiments, the ratio of occurrences (or copies or numbers) of targets to be labeled to different barcode sequences can be at least 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 1:30, 1:40, 1:50, 1:60, 1:70, 1:80, 1:90, or more. In some embodiments, each of the plurality of oligonucleotides further comprises a sample label, a universal label, or both. The particle can be, for example, a nanoparticle or a microparticle.

[0084] The size of the beads can vary. For example, the diameter of the beads can range from 0.1 micrometers to 50 micrometers. In some embodiments, the diameter of the beads can be 0.1, 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 micrometers, or a number or range between any two of these values, or can be approximately these values ​​or such numbers or ranges. The diameter of the beads may be related to the diameter of the wells of the substrate. In some embodiments, the diameter of the beads may be 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or a number or range between any two of these values, longer or shorter than the diameter of the wells, may be about these values ​​or such number or range, longer or shorter, may be at least these values ​​or such number or range, longer or shorter, or may be up to these values ​​or such number or range, longer or shorter. The diameter of the beads may be related to the diameter of a cell (e.g., a single cell surrounded by a well of the substrate). The diameter of the beads may be related to the diameter of a cell (e.g., a single cell surrounded by a well of the substrate). In some embodiments, the diameter of the bead may be 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 250%, 300%, or a number or range between any two of these values, longer or shorter, longer or shorter than, approximately, at least, or up to, the diameter of the cell.

[0085] The beads may be bound and / or embedded in a substrate. The beads may be bound and / or embedded in a gel, hydrogel, polymer, and / or matrix. The spatial location of the beads within the substrate (e.g., gel, matrix, scaffold, or polymer) can be identified using spatial labels present in a barcode on the beads, which can serve as a positional address. Examples of beads include, but are not limited to, streptavidin beads, agarose beads, magnetic beads, Dynabeads®, MACS® microbeads, antibody-conjugated beads (e.g., anti-immunoglobulin microbeads), Protein A-conjugated beads, Protein G-conjugated beads, Protein A / G-conjugated beads, Protein L-conjugated beads, oligo(dT)-conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, and BcMag™ carboxyl-terminated magnetic beads.

[0086] The beads can be associated with (e.g., impregnated with) quantum dots or fluorescent dyes to make them fluorescent in one or more optical channels. The beads can be associated with iron oxide or chromium oxide to make them paramagnetic or ferromagnetic. The beads can be identifiable. For example, the beads can be imaged using a camera. The beads can have a detectable code associated with them. For example, the beads can include a barcode. The beads can change size, for example, due to swelling in an organic or inorganic solution. The beads can be hydrophobic. The beads can be hydrophilic. The beads can be biocompatible. The solid support (e.g., beads) can be visualized. The solid support can include a visualization tag (e.g., a fluorescent dye). The solid support (e.g., beads) can be etched with an identifier (e.g., a number). The identifier can be visualized through imaging of the beads.

[0087] A solid support may comprise an insoluble, semi-soluble, or insoluble material. A solid support may be referred to as "functionalized" if it comprises a linker, scaffold, building block, or other reactive moiety attached thereto, whereas a solid support may be "non-functionalized" if it lacks such reactive moieties attached thereto. A solid support may be freely used in solution, e.g., in a microtiter well format, in a flow-through format, e.g., in a column, or in a dipstick.

[0088] The solid support may include a membrane, paper, plastic, coated surface, flat surface, glass, slide, chip, or any combination thereof. The solid support may take the form of a resin, gel, microsphere, or other geometric configuration. The solid support may include silica chips, microparticles, nanoparticles, plates, arrays, calipers, flat supports, such as glass fiber filters, glass surfaces, metal surfaces, metal surfaces (steel, gold silver, aluminum, silicone, and copper), glass supports, plastic supports, silicone supports, chips, filters, membranes, microwell plates, slides, multiwell plates, or plastic materials including membranes (e.g., made of polyethylene, polypropylene, polyamide, polyvinylidene difluoride), and / or wafers, combs, pins, or needles (e.g., arrays of pins suitable for combinatorial synthesis or analysis), or arrays of holes or nanoliter wells on a flat surface, such as wafers (e.g., silicone wafers), wafers with holes with or without filter bottoms. The solid support may include a polymer matrix (e.g., a gel, a hydrogel). The polymer matrix may be capable of penetrating intracellular spaces (e.g., around organelles). The polymer matrix may be capable of being pumped through the circulatory system.

[0089] Substrates and Microwell Arrays As used herein, a substrate may refer to a type of solid support. A substrate may refer to a solid support that may include a barcode or stochastic barcode of the present disclosure. A substrate may include, for example, a plurality of microwells. For example, a substrate may be a well array that includes two or more microwells. In some embodiments, a microwell may include a small reaction chamber with a defined volume. In some embodiments, a microwell may incorporate one or more cells. In some embodiments, a microwell may incorporate only one cell. In some embodiments, a microwell may incorporate one or more solid supports. In some embodiments, a microwell may incorporate only one solid support. In some embodiments, a microwell incorporates a single cell and a single solid support (e.g., a bead). A microwell may include a barcode reagent of the present disclosure.

[0090] Barcoding methods The present disclosure provides a method for estimating the number of distinct targets in distinct locations of a body sample (e.g., tissue, organ, tumor, cell). The method may include placing a barcode (e.g., stochastic barcode) in proximity to the sample, lysing the sample, associating distinct targets with the barcode, amplifying the targets, and / or digitally counting the targets. The method may further include analyzing and / or visualizing information obtained from the spatial labeling of the barcode. In some embodiments, the method includes visualizing a plurality of targets in the sample. Mapping the plurality of targets to a map of the sample may include creating a two-dimensional or three-dimensional map of the sample. The two-dimensional and three-dimensional maps may be created before or after barcoding (e.g., stochastic barcoding) the plurality of targets in the sample. Visualizing a plurality of targets in the sample may include mapping the plurality of targets to a map of the sample. Mapping the plurality of targets to a map of the sample may include creating a two-dimensional or three-dimensional map of the sample. The two-dimensional and three-dimensional maps may be generated before or after barcoding a plurality of targets in a sample. In some embodiments, the two-dimensional and three-dimensional maps may be generated before or after lysing the sample. Lysing the sample before or after generating the two-dimensional or three-dimensional map may include heating the sample, contacting the sample with a detergent, changing the pH of the sample, or any combination thereof.

[0091] In some embodiments, barcoding the multiple targets comprises hybridizing a plurality of barcodes to the multiple targets to generate barcoded targets (e.g., stochastically barcoded targets). Barcoding the multiple targets may comprise generating an indexed library of barcoded targets. Generating an indexed library of barcoded targets may be performed using a solid support comprising a plurality of barcodes (e.g., stochastic barcodes).

[0092] Contacting the sample with the barcode The present disclosure provides a method for contacting a sample (e.g., cells) with a substrate of the present disclosure. For example, a sample including a thin section of cells, an organ, or a tissue can be contacted with a barcode (e.g., a stochastic barcode). For example, the cells can be contacted by gravity flow, where the cells can settle and form a monolayer. The sample can be a tissue slice. The slice can be placed on a substrate. The sample can be one-dimensional (e.g., forming a planar surface). For example, the sample (e.g., cells) can be spread across the substrate by growing / culturing the cells on the substrate. When the barcode is in close proximity to the target, the target can hybridize to the barcode. The barcodes can be contacted in a non-depleting ratio so that each separate target can associate with a separate barcode of the present disclosure. To ensure efficient association between the target and the barcode, the target can be cross-linked to the barcode.

[0093] Cell lysis After partitioning of the cells and barcodes, the cells can be lysed to release the target molecules. Cell lysis can be achieved by any of a variety of means, for example, by chemical or biochemical means, by osmotic shock, or by thermal, mechanical, or optical lysis. Cells may be lysed by adding a cell lysis buffer containing a detergent (e.g., SDS, Li-dodecyl sulfate, Triton X-100, Tween-20, or NP-40), an organic solvent (e.g., methanol or acetone), or a digestive enzyme (e.g., proteinase K, pepsin, or trypsin), or any combination thereof. To increase the association of the target with the barcode, the diffusion rate of the target molecule may be altered, for example, by lowering the temperature and / or increasing the viscosity of the lysate.

[0094] In some embodiments, the sample may be lysed by using filter paper, which can be soaked with a lysis buffer over the filter paper, and pressure can be applied to the sample that can promote lysis of the sample and hybridization of the sample's targets to the substrate. In some embodiments, lysis can be performed by mechanical lysis, heat lysis, optical lysis, and / or chemical lysis. Chemical lysis may include the use of digestive enzymes such as proteinase K, pepsin, and trypsin. Lysis can be performed by the addition of a lysis buffer to the substrate. The lysis buffer may include Tris-HCl. The lysis buffer may include at least about 0.01, 0.05, 0.1, 0.5, or 1 M or more Tris-HCl. The lysis buffer may include up to about 0.01, 0.05, 0.1, 0.5, or 1 M or more Tris-HCl. The lysis buffer may include about 0.1 M Tris-HCl. The pH of the lysis buffer may be at least about or up to about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. In some embodiments, the pH of the lysis buffer is about 7.5. The lysis buffer may include a salt (e.g., LiCl). The salt concentration in the lysis buffer may be at least about or up to about 0.1, 0.5, or 1 M or higher. In some embodiments, the salt concentration in the lysis buffer is about 0.5 M. The lysis buffer may include a detergent (e.g., SDS, Li-dodecyl sulfate, triton X, tween, NP-40). The detergent concentration in the lysis buffer may be at least about, up to about, 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, or 7%, or higher. In some embodiments, the detergent concentration in the lysis buffer is about 1% Li-dodecyl sulfate. The time used in the method for lysis may depend on the amount of detergent used. In some embodiments, the more detergent used, the less time is required for lysis. The lysis buffer may include a chelating agent (e.g., EDTA, EGTA). The concentration of the chelating agent in the lysis buffer may be at least about or up to about 1, 5, 10, 15, 20, 25, or 30 mM. In some embodiments, the concentration of the chelating agent in the lysis buffer is about 10 mM. The lysis buffer may include a reducing agent (e.g., beta-mercaptoethanol, DTT).The concentration of the reducing reagent in the lysis buffer may be at least about or up to about 1, 5, 10, 15, or 20 mM. In some embodiments, the concentration of the reducing reagent in the lysis buffer is about 5 mM. In some embodiments, the lysis buffer may include about 0.1 M Tris-HCl (about pH 7.5), about 0.5 M LiCl, about 1% lithium dodecyl sulfate, about 10 mM EDTA, and about 5 mM DTT.

[0095] Lysing can be performed at a temperature of about 4, 10, 15, 20, 25, or 30° C. Lysing can be performed for about 1, 5, 10, 15, 20 minutes, or longer. Lysed cells may contain at least about 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, or more target nucleic acid molecules. Lysed cells may contain up to about 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, or more target nucleic acid molecules.

[0096] Binding of barcodes to target nucleic acid molecules After lysis of the cells and release of the nucleic acid molecules therefrom, the nucleic acid molecules may randomly associate with the barcodes of the co-localized solid support. The association may include hybridization of the target recognition region of the barcode to a complementary portion of the target nucleic acid molecule (e.g., the oligo(dT) of the barcode may interact with the poly(A) tail of the target). The assay conditions used for hybridization (e.g., buffer pH, ionic strength, temperature, etc.) may be selected to promote the formation of specific stable hybrids. In some embodiments, the nucleic acid molecules released from the lysed cells may associate with multiple probes on the substrate (e.g., hybridize to the probes on the substrate). If the probes include oligo(dT), the mRNA molecules may hybridize to the probes and be reverse transcribed. The oligo(dT) portion of the oligonucleotide may act as a primer for first strand synthesis of the cDNA molecules. For example, in the non-limiting example of barcoding shown in block 216 of FIG. 2, the mRNA molecules may hybridize to the barcodes on the beads. For example, the single-stranded nucleotide fragment may hybridize to the target binding region of the barcode.

[0097] The binding may further include ligating the target recognition region of the barcode with a portion of the target nucleic acid molecule. For example, the target binding region may include a nucleic acid sequence that may allow specific hybridization to a restriction site overhang (e.g., an EcoRI sticky end overhang). The assay procedure may further include treating the target nucleic acid with a restriction enzyme (e.g., EcoRI) to generate a restriction site overhang. The barcode may then be ligated to any nucleic acid molecule that includes a sequence complementary to the restriction site overhang. A ligase (e.g., T4 DNA ligase) may be used to connect the two fragments.

[0098] For example, in a non-limiting example of barcoding shown in block 220 of Figure 2, the labeled targets (e.g., target-barcode molecules) from multiple cells (or multiple samples) can then be pooled, for example, in a tube. The labeled targets can be pooled, for example, by collecting beads to which barcodes and / or target-barcode molecules are attached. Recovery of the solid support-based collection of bound target-barcode molecules can be achieved by the use of magnetic beads and an externally applied magnetic field. Once the target-barcode molecules are pooled, all further processing can proceed within a single reaction vessel. Further processing can include, for example, reverse transcription reactions, amplification reactions, cleavage reactions, dissociation reactions, and / or nucleic acid extension reactions. Further processing reactions can be performed within the microwells, i.e., without first pooling the labeled target nucleic acid molecules from multiple cells.

[0099] Reverse transcription The present disclosure provides methods for generating target-barcode conjugates using reverse transcription (e.g., block 224 of FIG. 2). The target-barcode conjugates may include a barcode and a complementary sequence of all or a portion of a target nucleic acid (i.e., a barcoded cDNA molecule, e.g., a stochastically barcoded cDNA molecule). Reverse transcription of the associated RNA molecule may occur by adding a reverse transcription primer along with a reverse transcriptase. The reverse transcription primer may be an oligo(dT) primer, a random hexanucleotide primer, or a target-specific oligonucleotide primer. The oligo(dT) primer may be 12-18 nucleotides in length, or may be about 12-18 nucleotides in length, and binds to an endogenous poly(A) tail at the 3' end of a mammalian mRNA. The random hexanucleotide primer may bind to the mRNA at various complementary sites. The target-specific oligonucleotide primer typically selectively primes the mRNA of interest.

[0100] In some embodiments, reverse transcription of the labeled RNA molecule can occur by addition of a reverse transcription primer. In some embodiments, the reverse transcription primer is an oligo(dT) primer, a random hexanucleotide primer, or a target-specific oligonucleotide primer. Generally, oligo(dT) primers are 12-18 nucleotides in length and bind to endogenous poly(A) tails at the 3' end of mammalian mRNAs. Random hexanucleotide primers can bind to mRNAs at various complementary sites. Target-specific oligonucleotide primers typically selectively prime the mRNA of interest. Reverse transcription can occur repeatedly to generate multiple labeled cDNA molecules. The methods disclosed herein can include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 reverse transcription reactions. The methods can include performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 reverse transcription reactions.

[0101] amplification One or more nucleic acid amplification reactions (e.g., block 228 of FIG. 2) can be performed to generate multiple copies of the labeled target nucleic acid molecule. Amplification can be performed in a multiplex manner, where multiple target nucleic acid sequences are amplified simultaneously. The amplification reaction can be used to add sequencing adapters to the nucleic acid molecule. The amplification reaction can include amplifying at least a portion of the sample label, if present. The amplification reaction can include amplifying at least a portion of the cell label and / or barcode sequence (e.g., molecular label). The amplification reaction can include amplifying at least a portion of the sample tag, cell label, spatial label, barcode sequence (e.g., molecular label), target nucleic acid, or a combination thereof. The amplification reaction may include amplifying 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 100%, or a range or number between any two of these values ​​of the plurality of nucleic acids. The method may further include performing one or more cDNA synthesis reactions to generate one or more cDNA copies of the target-barcode molecule that includes the sample label, cell label, spatial label, and / or barcode sequence (e.g., molecular label).

[0102] In some embodiments, amplification can be performed using polymerase chain reaction (PCR). As used herein, PCR can refer to a reaction for amplifying specific DNA sequences in vitro by simultaneous primer extension of complementary strands of DNA. As used herein, PCR can encompass derivatives of the reaction, including but not limited to RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, and assembly PCR.

[0103] Amplification of the labeled nucleic acid may include non-PCR-based methods. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or circle-circle amplification. Other non-PCR-based amplification methods include DNA-dependent RNA polymerase-driven RNA transcription amplification or multiple cycles of RNA-directed DNA synthesis and transcription to amplify DNA or RNA targets, ligase chain reaction (LCR), and Qβ replicase (Qβ) method, the use of palindromic probes, strand displacement amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods in which a primer is hybridized to a nucleic acid sequence and the resulting double strand is cleaved before extension reaction and amplification, strand displacement amplification using a nucleic acid polymerase lacking 5' exonuclease activity, rolling circle amplification, and branched extension amplification (RAM). In some embodiments, the amplification does not generate circularized transcripts.

[0104] In some embodiments, the methods disclosed herein further include performing a polymerase chain reaction on the labeled nucleic acid (e.g., labeled RNA, labeled DNA, labeled cDNA) to generate a labeled amplicon (e.g., a stochastically labeled amplicon). The labeled amplicon may be a double-stranded molecule. The double-stranded molecule may include a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized to a DNA molecule. One or both strands of the double-stranded molecule may include a sample label, a spatial label, a cell label, and / or a barcode sequence (e.g., a molecular label). The labeled amplicon may be a single-stranded molecule. The single-stranded molecule may include DNA, RNA, or a combination thereof. The nucleic acid of the present disclosure may include a synthetic or modified nucleic acid.

[0105] Amplification may include the use of one or more non-natural nucleotides. Non-natural nucleotides may include photolabile or triggering nucleotides. Examples of non-natural nucleotides include, but are not limited to, peptide nucleic acid (PNA), morpholino nucleic acid and locked nucleic acid (LNA), as well as glycol nucleic acid (GNA) and threose nucleic acid (TNA). Non-natural nucleotides may be added to one or more cycles of the amplification reaction. Addition of non-natural nucleotides may be used to identify products as specific cycles or time points of the amplification reaction.

[0106] Conducting one or more amplification reactions may include the use of one or more primers. The one or more primers may, for example, comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 or more nucleotides. The one or more primers may comprise at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 or more nucleotides. The one or more primers may comprise less than 12 to 15 nucleotides. The one or more primers may anneal to at least a portion of the multiple labeled targets (e.g., stochastically labeled targets). The one or more primers may anneal to the 3' or 5' ends of the multiple labeled targets. The one or more primers may anneal to an internal region of the multiple labeled targets. The internal region may be at least about 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of the multiple labeled targets. The one or more primers may comprise a constant panel of primers. The one or more primers may include at least one or more custom primers. The one or more primers may include at least one or more control primers. The one or more primers may include at least one or more gene-specific primers.

[0107] The one or more primers may include a universal primer. The universal primer may anneal to the universal primer binding site. The one or more custom primers may anneal to a first sample label, a second sample label, a spatial label, a cell label, a barcode sequence (e.g., a molecular label), a target, or any combination thereof. The one or more primers may include a universal primer and a custom primer. The custom primers may be designed to amplify one or more targets. The targets may include a subset of all nucleic acids in one or more samples. The targets may include a subset of all labeled targets in one or more samples. The one or more primers may include at least 96 or more custom primers. The one or more primers may include at least 960 or more custom primers. The one or more primers may include at least 9600 or more custom primers. The one or more custom primers may anneal to two or more different labeled nucleic acids. The two or more different labeled nucleic acids may correspond to one or more genes.

[0108] Any amplification scheme can be used in the disclosed method. For example, in one scheme, the first PCR can amplify the molecules bound to the beads using a gene-specific primer and a primer for the sequence of universal Illumina sequencing primer 1. The second PCR can amplify the first PCR product using a nested gene-specific primer adjacent to the sequence of Illumina sequencing primer 2 and a primer for the sequence of universal Illumina sequencing primer 1. The third PCR adds P5 and P7 and a sample index to place the PCR product into an Illumina sequencing library. Sequencing using 150bp x 2 sequencing can reveal cell label and barcode sequences (e.g., molecular labels) on read 1, genes on read 2, and sample index on index 1 read.

[0109] In some embodiments, chemical cleavage can be used to remove nucleic acids from a substrate. For example, chemical groups or modified bases present in the nucleic acid can be used to facilitate its removal from the solid support. For example, an enzyme can be used to remove nucleic acids from a substrate. For example, a nucleic acid can be removed from a substrate by restriction endonuclease digestion. For example, treatment of a nucleic acid containing dUTP or ddUTP with uracil-d-glycosylase (UDG) can be used to remove nucleic acids from a substrate. For example, an enzyme that performs nucleotide excision, e.g., a base excision repair enzyme, e.g., an apurinic / apyrimidinic (AP) endonuclease, can be used to remove nucleic acids from a substrate. In some embodiments, a photocleavable group and light can be used to remove nucleic acids from a substrate. In some embodiments, a cleavable linker can be used to remove nucleic acids from a substrate. For example, the cleavable linker can include at least one of biotin / avidin, biotin / streptavidin, biotin / neutravidin, Ig-Protein A, a photolabile linker, an acid or base labile linker group, or an aptamer.

[0110] If the probe is gene specific, the molecule may be hybridized to the probe and reverse transcribed and / or amplified. In some embodiments, the nucleic acid may be amplified after it is synthesized (e.g., reverse transcribed). Amplification may be performed in a multiplex manner, where multiple target nucleic acid sequences are amplified simultaneously. Amplification may add sequencing adapters to the nucleic acid.

[0111] In some embodiments, amplification can be performed on the substrate, for example, using bridge amplification. To generate ends compatible with bridge amplification using oligo(dT) probes on the substrate, a homopolymeric tail can be added to the cDNA. In bridge amplification, the primer complementary to the 3' end of the template nucleic acid can be the first primer of each pair covalently attached to the solid particle. When a sample containing the template nucleic acid is contacted with the particle and one thermal cycle is performed, the template molecule can be annealed to the first primer, and the first primer can be extended in the forward direction by adding nucleotides to form a double-stranded molecule consisting of the template molecule and a newly formed DNA strand complementary to the template. In the heating step of the next cycle, the double-stranded molecule can be denatured, releasing the template molecule from the particle and leaving the complementary DNA strand attached to the particle through the first primer. In the annealing stage of the subsequent annealing and extension step, the complementary strand can hybridize to a second primer complementary to the segment of the complementary strand at the position removed from the first primer. By this hybridization, the complementary strand can form a bridge between the first and second primers, immobilized by covalent bond to the first primer and by hybridization to the second primer. In the extension step, the second primer can be extended in the opposite direction by adding nucleotides in the same reaction mixture, thereby converting the bridge into a double-stranded bridge. The next cycle then begins, and the double-stranded bridge is denatured to obtain two single-stranded nucleic acid molecules, one end of which is bound to the particle surface through the first and second primers, respectively, and the other end of which is unbound, respectively. In the annealing and extension step of this second cycle, each strand can hybridize to a further previously unused complementary primer on the same particle to form a new single-stranded bridge. The two previously unused primers hybridized at this point extend to convert the two new bridges into double-stranded bridges.

[0112] The amplification reaction may include amplifying at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 100% of the plurality of nucleic acids. The amplification of the labeled nucleic acid may include PCR-based or non-PCR-based methods. The amplification of the labeled nucleic acid may include exponential amplification of the labeled nucleic acid. The amplification of the labeled nucleic acid may include linear amplification of the labeled nucleic acid. The amplification may be performed by polymerase chain reaction (PCR). PCR may refer to a reaction for in vitro amplification of specific DNA sequences by simultaneous primer extension of complementary strands of DNA. PCR may encompass derivatives of the reaction, including, but not limited to, RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, suppression PCR, semi-suppressive PCR, and assembly PCR.

[0113] In some embodiments, the amplification of the labeled nucleic acid comprises a non-PCR-based method. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or circle-circle amplification. Other non-PCR-based amplification methods include DNA-dependent RNA polymerase-driven RNA transcription amplification or multiple cycles of RNA-directed DNA synthesis and transcription to amplify DNA or RNA targets, ligase chain reaction (LCR), Qβ replicase (Qβ) method, use of palindromic probes, strand displacement amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods in which a primer is hybridized to a nucleic acid sequence and the resulting duplex is cleaved before extension reaction and amplification, strand displacement amplification using a nucleic acid polymerase lacking 5' exonuclease activity, rolling circle amplification, and / or branched extension amplification (RAM).

[0114] In some embodiments, the methods disclosed herein further comprise performing a nested polymerase chain reaction on the amplified amplicon (e.g., target). The amplicon may be a double-stranded molecule. The double-stranded molecule may comprise a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized to a DNA molecule. One or both strands of the double-stranded molecule may comprise a sample tag or a molecular identifier label. Alternatively, the amplicon may be a single-stranded molecule. The single-stranded molecule may comprise DNA, RNA, or a combination thereof. The nucleic acid of the present invention may comprise a synthetic or modified nucleic acid. In some embodiments, the methods include repeatedly amplifying the labeled nucleic acid to generate multiple amplicons. The methods disclosed herein may include performing amplification reactions about, or at least about, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 times, or a number or range between any two of these values.

[0115] The amplification may further include adding one or more control nucleic acids to the one or more samples comprising the plurality of nucleic acids. The amplification may further include adding one or more control nucleic acids to the plurality of nucleic acids. The control nucleic acids may include a control label. Amplification may include the use of one or more non-natural nucleotides. Non-natural nucleotides may include photolabile and / or trigger nucleotides. Examples of non-natural nucleotides include, but are not limited to, peptide nucleic acid (PNA), morpholino nucleic acid and locked nucleic acid (LNA), as well as glycol nucleic acid (GNA) and threose nucleic acid (TNA). Non-natural nucleotides may be added to one or more cycles of the amplification reaction. The addition of non-natural nucleotides may be used to identify products as specific cycles or time points of the amplification reaction.

[0116] The amplification reaction or reactions may include the use of one or more primers. The one or more primers may include one or more oligonucleotides. The one or more oligonucleotides may include at least about 7-9 nucleotides. The one or more oligonucleotides may include less than 12-15 nucleotides. The one or more primers may anneal to at least a portion of the plurality of labeled nucleic acids. The one or more primers may anneal to the 3' and / or 5' ends of the plurality of labeled nucleic acids. The one or more primers may anneal to an internal region of the plurality of labeled nucleic acids. The internal region may be at least about 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of the plurality of labeled nucleic acids. The one or more primers may comprise a fixed panel of primers. The one or more primers may include at least one or more custom primers. The one or more primers may include at least one or more control primers. The one or more primers may include at least one or more housekeeping gene primers. The one or more primers may include a universal primer. The universal primer may anneal to a universal primer binding site. The one or more custom primers may anneal to a first sample tag, a second sample tag, a molecular identifier label, a nucleic acid or a product thereof. The one or more primers may include a universal primer and a custom primer. The custom primer may be designed to amplify one or more target nucleic acids. The target nucleic acid may include a subset of the total nucleic acids in one or more samples. In some embodiments, the primer is a probe bound to the array of the present disclosure.

[0117] In some embodiments, barcoding (e.g., stochastically barcoding) a plurality of targets in a sample further comprises generating an indexed library of barcoded targets (e.g., stochastically barcoded targets) or barcoded fragments of those targets. The barcode sequences of the different barcodes (e.g., molecular labels of the different stochastic barcodes) may be different from each other. Generating an indexed library of barcoded targets comprises generating a plurality of indexed polynucleotides from the plurality of targets in the sample. For example, for an indexed library of barcoded targets including a first indexed target and a second indexed target, the label region of the first indexed polynucleotide may differ from the label region of the second indexed polynucleotide by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 nucleotides, or a number or range between any two of these values, or by about these values ​​or such number or range, or by at least these values ​​or such number or range, or by up to these values ​​or such number or range. In some embodiments, generating an indexed library of barcoded targets includes contacting a plurality of targets, e.g., mRNA molecules, with a plurality of oligonucleotides comprising a poly(T) region and a label region, and performing first strand synthesis using a reverse transcriptase to generate single-stranded labeled cDNA molecules, each of which comprises a cDNA region and a label region, where the plurality of targets comprises at least two mRNA molecules of different sequences, and the plurality of oligonucleotides comprises at least two oligonucleotides of different sequences. Generating an indexed library of barcoded targets may further include amplifying the single-stranded labeled cDNA molecules to generate double-stranded labeled cDNA molecules, and performing nested PCR on the double-stranded labeled cDNA molecules to generate labeled amplicons. In some embodiments, the method may include generating adapter-labeled amplicons.

[0118] Barcoding (e.g., stochastic barcoding) may include labeling individual nucleic acid (e.g., DNA or RNA) molecules with nucleic acid barcodes or tags. In some embodiments, this includes adding DNA barcodes or tags to cDNA molecules as they are generated from mRNA. Nested PCR can perform PCR amplification bias minimization. For example, adapters can be added for sequencing using NGS. Sequencing results can be used, for example, to determine the sequence of cellular labels, molecular labels, and nucleotide fragments of one or more copies of the target of block 232 of FIG. 2.

[0119] 3 is a schematic diagram illustrating a non-limiting exemplary process for generating an indexed library of barcoded targets (e.g., stochastically barcoded targets), e.g., barcoded mRNAs or fragments thereof. As shown in step 1, a reverse transcription process can encode each mRNA molecule with a unique molecular label sequence, a cellular label sequence, and a universal PCR site. In particular, the RNA molecule 302 can be reverse transcribed to generate labeled cDNA molecules 304 that include a cDNA region 306 by hybridization (e.g., stochastic hybridization) of a set of barcodes (e.g., stochastic barcodes) 310 to a poly(A) tail region 308 of the RNA molecule 302. Each of the barcodes 310 can include a target binding region, e.g., a poly(dT) region 312, a label region 314 (e.g., a barcode sequence or molecule), and a universal PCR region 316.

[0120] In some embodiments, the cell label sequence may comprise 3-20 nucleotides. In some embodiments, the molecular label sequence may comprise 3-20 nucleotides. In some embodiments, each of the plurality of stochastic barcodes further comprises one or more of a universal label and a cell label, wherein the universal label is the same for the plurality of stochastic barcodes on the solid support, and the cell label is the same for the plurality of stochastic barcodes on the solid support. In some embodiments, the universal label may comprise 3-20 nucleotides. In some embodiments, the cell label comprises 3-20 nucleotides.

[0121] In some embodiments, the label region 314 may include a barcode sequence or molecular label 318 and a cell label 320. In some embodiments, the label region 314 may include one or more of a universal label, a dimensional label, and a cell label. The barcode sequence or molecular label 318 may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range of nucleotides in between any of these values, or may be approximately these values ​​or such number or range of nucleotides in length, or may be at least these values ​​or such number or range of nucleotides in length, or may be up to these values ​​or such number or range of nucleotides in length. A cell label 320 may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range of nucleotides in between any of these values, or may be approximately, or may be at least, or may be up to, any of these values ​​or such numbers or ranges of nucleotides in length. A universal label may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range of nucleotides in between any of these values, or may be approximately, or may be at least, or may be up to, any of these values ​​or such numbers or ranges of nucleotides in length. The universal label may be the same for multiple stochastic barcodes on a solid support, and the cell label may be the same for multiple stochastic barcodes on a solid support.A dimension marker may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range between any of these values, in length, or may be approximately these values ​​or such number or range of nucleotides in length, or may be at least these values ​​or such number or range of nucleotides in length, or may be up to these values ​​or such number or range of nucleotides in length.

[0122] In some embodiments, label region 314 may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or a number or range of any of these values ​​between, or include about, or include at least, or include at most, or include at most, or include at most, Each label may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range between any of these values, in length, or may be approximately, at least, or up to these values ​​or such number or range of nucleotides in length. A set of barcodes or stochastic barcodes 310 may be 10, 20, 40, 50, 70, 80, 90, 100, or a number or range between any of these values, in length, or may be approximately, at least, or up to these values ​​or such number or range of nucleotides in length. 2 , 10 3 , 10 4 , 10 5 , 10 6 , 107 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 , or between any of these values, or about these values, or such number or range of barcodes or stochastic barcodes 310, or at least these values, or such number or range of barcodes or stochastic barcodes 310, or up to these values, or such number or range of barcodes or stochastic barcodes 310. The set of barcodes or stochastic barcodes 310 may also each contain a unique labeled region 314, for example. The labeled cDNA molecules 304 may be purified to remove excess barcodes or stochastic barcodes 310. Purification may include Ampure bead purification.

[0123] As shown in step 2, the products from the reverse transcription process in step 1 can be pooled in one tube and PCR amplified using a first PCR primer pool and a first universal PCR primer. Pooling is possible due to the unique label region 314. In particular, the labeled cDNA molecules 304 can be amplified to generate nested PCR labeled amplicons 322. The amplification can include multiplex PCR amplification. The amplification can include multiplex PCR amplification using 96 multiplex primers in a single reaction volume. In some embodiments, the multiplex PCR amplification can include multiplex PCR amplification using 10, 20, 40, 50, 70, 80, 90, ... 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 1013 , 10 14 , 10 15 , 10 20 , or any of these values, or between these values, or about these values, or at least these values, or at most these values. The amplification may include a first PCR primer pool 324 including custom primers 326A-C targeting specific genes and a universal primer 328. The custom primer 326 may hybridize to a region within the cDNA portion 306' of the labeled cDNA molecule 304. The universal primer 328 may hybridize to the universal PCR region 316 of the labeled cDNA molecule 304.

[0124] As shown in step 3 of FIG. 3, the products from the PCR amplification in step 2 may be amplified using a nested PCR primer pool and a second universal PCR primer. Nested PCR can minimize PCR amplification bias. In particular, the nested PCR labeled amplicons 322 may be further amplified by nested PCR. Nested PCR may include multiplex PCR including a nested PCR primer pool 330 of nested PCR primers 332a-c and a second universal PCR primer 328' in a single reaction volume. The nested PCR primer pool 328 may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or a number or range of different nested PCR primers 330 between any of these values, or may contain about these values ​​or such number or range of different nested PCR primers 330, or may contain at least these values ​​or such number or range of different nested PCR primers 330, or may contain up to these values ​​or such number or range of different nested PCR primers 330. The nested PCR primers 332 may contain an adaptor 334 and hybridize to a region within the cDNA portion 306'' of the labeled amplicon 322. Universal primer 328' contains adaptor 336 and can hybridize to universal PCR region 316 of labeled amplicon 322. Thus, step 3 generates adaptor-labeled amplicon 338. In some embodiments, nested PCR primer 332 and second universal PCR primer 328' may not contain adaptors 334 and 336. Instead, adaptors 334 and 336 can ligate to the product of nested PCR to generate adaptor-labeled amplicon 338.

[0125] As shown in step 4, the PCR products from step 3 can be PCR amplified for sequencing using library amplification primers. In particular, adapters 334 and 336 can be used to perform one or more additional assays on adapter-labeled amplicons 338. Adapters 334 and 336 can be hybridized with primers 340 and 342. One or more primers 340 and 342 can be PCR amplification primers. One or more primers 340 and 342 can be sequencing primers. One or more adapters 334 and 336 can be used for further amplification of adapter-labeled amplicons 338. One or more adapters 334 and 336 can be used for sequencing of adapter-labeled amplicons 338. Primer 342 can contain a plate index 344, which allows amplicons generated using the same set of barcodes or stochastic barcodes 310 to be sequenced in a single sequencing reaction using NGS.

[0126] Full-length single-cell RNA sequencing High-throughput single-cell RNA sequencing has changed the understanding of complex and heterogeneous biological samples. However, most methods allow only 3' analysis of mRNA transcript information, which may limit the analysis of highly variable loci due to rearrangements such as splice variants, alternative transcription start sites, and VDJ junctions of T-cell and B-cell receptors and antibodies. For both T-cells and B-cells, currently available C-priming-based approaches can read into the V(D)J but lose the upstream V region. Thus, currently available methods may limit the ability to obtain information on full-length nucleic acid targets (e.g., transcripts containing V(D)J). A particular problem in the art is the need to know the VDJ sequences as longer reads, since there are many VDJs due to the large number of possible rearrangement events. There is a need for methods to both count sequences (e.g., transcripts containing V(D)J) and identify said sequences (especially full-length sequences).

[0127] In some embodiments, methods are provided for obtaining full-length V(D)J information (e.g., by Illumina sequencing on the Rhapsody system). T and B cell receptors contain V segments, D segments (for TCR beta and BCR heavy chains only), J segments, and constant regions at the 3' end of the mRNA. The CDR3s composed of the V(D)J junctions contain the bulk of the repertoire diversity and are short enough to be sequenced on the Illumina short read platform. However, full-length V and D and J segment information is also useful and cannot be readily obtained without long-read sequencing technology, as the performance of Illumina short reads limits the ability to obtain full-length V(D)J information. The methods provided herein allow users to obtain both CDR3 information and full-length V segments, full-length D segment sequences, and / or full-length J segment sequences from a single library and corresponding sequencing run on an Illumina sequencer. Thus, some embodiments of the methods provided herein obtain full-length immune receptor mRNA sequences.

[0128] The disclosure herein includes a method for generating a whole transcriptome analysis (WTA) library from a DNA product (e.g., a DNA product of a BD Rhapsody Single Cell Analysis System) for sequencing on a corresponding sequencer (e.g., an Illumina® sequencer). In some embodiments, targeted RNA analysis generally provides higher sensitivity for low-expression targets, while WTA RNA analysis provides a broader range of genes to be examined. In some embodiments, successful amplification of a desired target may require a thorough understanding of the 3' end of each RNA and the use of the polyadenylation site at the 3' end of each RNA in the model system of interest. The WTA method, when used, for example, with BD Rhapsody, allows for several levels of information to be obtained, including: 1) identifying interesting RNA targets in those model systems to identify a panel of genes for design; 2) characterizing the 3' end of the transcriptome in those model systems as input into a new generation of PCR panel design for BD Rhapsody; and 3) allowing users to make biological discoveries by assaying a broader range of genes compared to standard BD Rhapsody targeting approaches.

[0129] In some embodiments, systems, methods, compositions, and kits are provided for 3'-based, internal-based, and / or 5'-based whole transcriptome analysis (WTA). The disclosure herein includes methods and compositions for 5', 3', and internal WTA library generation. The disclosed methods and compositions can be used for full-length single-cell RNA sequencing, for example using the Rhapsody system. The disclosed methods and compositions allow for full-length whole transcriptome analysis (WTA) sequencing of mRNA in single cells using the BD Rhapsody System. Currently available WTA assays sequence the 3' end of the mRNA by capturing the polyA tail at the 3' end of the mRNA and adding a unique molecular index (UMI) and a cellular label (CL). Additionally, currently available methods can profile the 5' end of the mRNA by adding a sequence to the 5' end of the cDNA that also crosslinks to the bead and adding a UMI and CL to the 5' end of the transcript. However, only small fragments (usually about 300-600 base pairs) can be profiled by sequencers. This means that any region more than about 750 bases from the 3' or 5' end cannot be profiled because there are no available methods for adding UMIs and CLs to internal regions. Currently available high-throughput scRNAseq methods provide sequencing of the 3' or 5' end of a transcript. Some currently available methods allow for resolution of either 3' or 5', but not both simultaneously. Full-length RNA is not captured by other methods. Currently available low-throughput scRNAseq methods may provide full-length sequencing, but at a much lower maximum cell number and are a highly labor-intensive process. Since Illumina devices are limited in sequencing long sequences, new methods and compositions are needed. The methods and compositions provided herein can overcome these disadvantages of currently available methods. In some embodiments, a full-length scRNAseq solution for high-throughput experiments is provided.The systems, methods, compositions, and kits provided herein can, in some embodiments, be used in conjunction with the methods and compositions described in PCT Patent Application No. PCT / US22 / 76366, entitled "FULL LENGTH SINGLE CELL RNA SEQUENCING," filed September 13, 2022, the contents of which are incorporated herein by reference in their entirety.

[0130] In some embodiments, compositions and methods are provided for performing full-length RNA sequencing (e.g., using the Rhapsody system). Currently available workflows provide 3' or 5' mRNA sequencing analysis at the single cell level, but not full length. The currently available methods cannot provide internal sequences of RNA because 1) the size of libraries that can be sequenced on an Illumina sequencer is limited, and 2) to obtain single cell information, cell labels and UMIs (e.g., molecular labels) should be present at the ends of the library. Current capture sequences for RNA / cDNA are through polyA (3' end of mRNA) or template switch oligonucleotides (5' end of mRNA). Thus, these approaches provide only 3' and 5' end sequences. To link cell labels to internal sequences, it is important to place the oligonucleotides near the internal region. In some embodiments of the methods and compositions provided herein, randomer / internal seq of targeted panel sequences on beads is used to add cell label information associated with internal mRNA sequences. In some embodiments, the approach described herein utilizes a polymerase with strand displacement activity to generate multiple copies of the internal sequence, and / or a polymerase without strand displacement activity to generate fewer regions of the internal sequence. These two approaches can be used depending on the purpose of full-length sequencing. If the user desires full coverage cDNA sequence, the first approach can be used. If the user desires sequencing of the internal region, but does not desire full coverage due to the high cost of sequencing, the latter method is suitable. If the user desires the internal sequence of all mRNAs, a randomer can be used. If the user desires a specific set of targeted genes (i.e., mutation hotspots) as internal sequences, a targeted gene-specific sequence can be added.4A-4H show schematic diagrams of non-limiting exemplary workflows for determining the sequence of a nucleic acid target (e.g., the V(D)J region of an immune receptor) and full-length whole transcriptome analysis (WTA) using the barcoding methods and compositions provided herein. The workflows can include one or more of the following steps: reverse transcription 400a, denaturation 400b, intermolecular hybridization 400c, contact with a cleavage agent 400d, contact with a polymerase without strand displacement activity 400e, contact with a polymerase with strand displacement activity 400f, denaturation and primer hybridization 400g, denaturation and primer hybridization 400h, primer extension, library generation, and sequencing 400i, and primer extension, library generation, and sequencing 400j. The nucleic acid target 406 (e.g., mRNA) can include a coding sequence (e.g., 414r) and a polyA tail (e.g., 408). The first plurality of oligonucleotide barcodes 402a may include a first universal sequence 426, a first molecular label 422, and a sequence complementary to at least a portion of a nucleic acid target (e.g., first target binding region 404). Each oligonucleotide barcode 416 (e.g., 416a1, 416a2) of the second plurality of oligonucleotide barcodes may include a second universal sequence 428, a cleavage domain 418, and a blocking group 420, a second molecular label 430, and a second target binding region (403). The cleavage domain may be located 5' to the blocking group. The oligonucleotide barcode may include a cell label 432. The oligonucleotide barcode may be attached to a solid support 401. The workflow may include production of one or more of the following products: 402b, 402c, 414c, 402c, 416c, 416d, and 416e. The product may include an element reverse complement (rc) as described herein. The workflow may include contacting the product with a random primer 446, which may include a third universal sequence 448.

[0131] In some embodiments, the solid supports provided herein (e.g., Rhapsody beads) have oligo dT capture sequences and / or template switch oligo capture sequences on the beads for capture through the 3' or 5' end of mRNA / cDNA. The addition of randomer / gene specific sequences to capture oligonucleotides on the solid support (e.g., Rhapsody beads) has not been used. However, as described herein, the presence of these oligos on the solid support (e.g., Rhapsody beads) may allow for the generation of libraries containing internal sequences to enable full-length RNAseq. The addition of blocked primers allows the user to control the time of primer use, which may reduce noise generation during library preparation.

[0132] 5A-5E show schematic diagrams of non-limiting exemplary workflows for determining the sequence of a nucleic acid target (e.g., the V(D)J region of an immune receptor) and full-length whole transcriptome analysis (WTA) using the barcoding methods and compositions provided herein.

[0133] FIG. 5A shows a non-limiting exemplary solid support (e.g., beads). Full-length RNAseq can include the beads (e.g., rhapsody beads) and associated oligonucleotides shown. The workflow can include mRNA capture and reverse transcription, followed by mRNA denaturation or use of RNase H to remove the RNA (top oligonucleotide). A randomer (UMI, molecular beacon) or gene-specific sequence can have a rhAmp tail to prevent the sequence from binding to the RNA and reverse transcription (bottom oligonucleotide). rhAmp PCR can include annealing of a blocked inactive primer to the target. The primer can have an overhang and an internal RNA base 5' to the blocking moiety. RNase H2 can recognize the hybridized internal RNA base, and cleavage can occur 5' to the RNA base upon hybridization of the primer to the target DNA. DNA polymerase can then extend the newly deblocked primer.

[0134] FIG. 5B shows non-limiting exemplary embodiments of the first and second workflows provided herein. In some embodiments, the workflow is a first workflow that includes full-length RNAseq without strand-displacement activity. The first workflow may include hybridization of randomers on cDNA. The first workflow may include RNase H2 treatment to release blockers and polymerase reaction. In some embodiments of the first workflow, polymerization stops at the end of other randomer binding regions due to the lack of strand-displacement activity. This first workflow allows for the generation of DNA pieces in different regions of cDNA. This first workflow may be ideal when a user desires a medium level of full-length coverage. In some embodiments, the workflow is a second workflow that includes full-length RNAseq with strand-displacement activity. The second workflow may include hybridization of randomers on cDNA. The second workflow may include RNase H2 treatment to release blockers and polymerase reaction. In some embodiments, in the second workflow, polymerization continues to proceed due to strand displacement activity, and a large number of different sizes of DNA are generated. Figure 5C shows a non-limiting exemplary embodiment of the second workflow (full-length RNAseq (with strand displacement activity)) provided herein. In the second workflow, polymerization continues to proceed due to strand displacement activity, and a large number of different sizes of DNA are generated. This is ideal when user wants to completely cover full-length sequences.

[0135] FIG. 5D shows a non-limiting exemplary embodiment of full-length RNAseq without PTA. In some embodiments, polymerization produces products of different sizes. FIG. 5E shows a non-limiting exemplary embodiment of RPE-PCR (using random with R2 sequence). UMIs from T1 oligos can be used as randomers and can be used to further filter PCR amplicons. In some embodiments, UMIs from T1 oligos can be used to count the number of molecules. In some embodiments, UMIs from T1 oligos are not used to count the number of molecules (e.g., there are more than one UMI / mRNA transcript). In some embodiments, RPE is repeated to have R2 adapters. In some embodiments, randomers from other beads can be attached to cDNA. In some embodiments, instead of using randomers, sequence-specific primers can be used to generate complete coverage of a particular gene of interest (e.g., cancer-associated mutation detection in gene candidates) at the single-cell level.

[0136] The disclosure herein includes a method for labeling a nucleic acid target in a sample. In some embodiments, the method includes contacting a copy of the nucleic acid target with a first plurality of oligonucleotide barcodes, each oligonucleotide barcode of the first plurality of oligonucleotide barcodes comprising a first universal sequence, a first molecular label, and a first target binding region capable of hybridizing to the nucleic acid target. In some embodiments, the method includes extending the first plurality of oligonucleotide barcodes hybridized to the copy of the nucleic acid target to generate a plurality of barcoded nucleic acid molecules, each of which comprises the first universal sequence, the first molecular label, and a sequence complementary to at least a portion of the nucleic acid target. In some embodiments, the method includes contacting the barcoded nucleic acid molecule with a second plurality of oligonucleotide barcodes for hybridization. In some embodiments, each oligonucleotide barcode of the second plurality of oligonucleotide barcodes comprises a second universal sequence, a cleavage domain, and a blocking group. In some embodiments, a blocking group can prevent extension of the oligonucleotide barcode and a cleavage domain is located 5' to the blocking group, such that when the cleavage domain hybridizes to the barcoded nucleic acid molecule, the oligonucleotide barcode can be cleaved by a cleavage enzyme at a point within or adjacent to the cleavage domain. In some embodiments, the method includes contacting a second plurality of oligonucleotide barcodes hybridized to the barcoded nucleic acid molecule with a cleavage enzyme, thereby removing the blocking group from the oligonucleotide barcodes. In some embodiments, the method includes extending a 3' end of an oligonucleotide barcode of the second plurality of oligonucleotide barcodes hybridized to the barcoded nucleic acid molecule to generate a plurality of extended barcoded nucleic acid molecules.

[0137] In some embodiments, each extended barcoded nucleic acid molecule of the plurality of extended barcoded nucleic acid molecules comprises at least a portion of the sequence of the nucleic acid target. In some embodiments, the cleavage enzyme is a ribonuclease H enzyme and / or a ribonuclease H2 enzyme, and the ribonuclease H2 enzyme may be a Pyrococcus abysilibonuclease H2 enzyme. In some embodiments, the cleavage enzyme is a hot-start cleavage enzyme that is thermostable and has reduced activity at low temperatures. In some embodiments, the hot-start cleavage enzyme is a Pyrococcus abysilibonuclease H2 that includes: (a) a G12A amino acid substitution; (b) a P13T amino acid substitution; (c) a G169A amino acid substitution; or (d) a combination thereof. In some embodiments, the cleavage enzyme is chemically modified. In some embodiments, the cleavage enzyme is a chemically modified hot-start cleavage enzyme that is thermostable and has reduced activity at low temperatures, and the cleavage enzyme may be reversibly inactivated by interaction with an antibody at low temperatures. In some embodiments, the cleavage domain comprises one or more ribonucleotides capable of being cleaved by a RNase H enzyme. In some embodiments, the cleavage domain comprises one or more of the following moieties: a DNA residue, an abasic residue, a modified nucleoside, or a modified phosphate internucleotide linkage. In some embodiments, the cleavage domain comprises at least one RNA base. In some embodiments, the cleavage domain comprises one or more 2'-modified nucleosides, where the one or more modified nucleosides may be 2'-fluoro nucleosides. In some embodiments, the blocking group is attached to the 3' terminal nucleotide of the oligonucleotide barcode. In some embodiments, the blocking group is at or near the 3' end of the oligonucleotide barcode. In some embodiments, the blocking group is a 2',3'-dideoxynucleotide, a ribonucleotide residue, a 2',3'-SH nucleotide, or a 2'-O-PO3 nucleotide. In some embodiments, the blocking group comprises a non-nucleotide modification. In some embodiments, the blocking group further comprises a naphthyl-azo compound, a spacer, and / or biotin.

[0138] In some embodiments, extending the 3' ends of the oligonucleotide barcodes of the second plurality of oligonucleotide barcodes comprises extending the 3' ends of the oligonucleotide barcodes of the second plurality of oligonucleotide barcodes with a DNA polymerase having strand displacement activity. In some embodiments, extending the 3' ends of the oligonucleotide barcodes of the second plurality of oligonucleotide barcodes with a DNA polymerase having strand displacement activity can generate an extended barcoded nucleic acid molecule comprising a complement of the first molecular label and a complement of the first universal sequence. In some embodiments, extending the 3' ends of the oligonucleotide barcodes of the second plurality of oligonucleotide barcodes comprises extending the 3' ends of the oligonucleotide barcodes of the second plurality of oligonucleotide barcodes with a DNA polymerase not having strand displacement activity.In some embodiments, the polymerase is selected from the group consisting of Phi29 DNA polymerase, E. coli DNA polymerase I, Bsu DNA polymerase, Bst DNA polymerase, Taq DNA polymerase, VENT™ DNA polymerase, DEEPVENT™ DNA polymerase, LongAmp® Taq DNA polymerase, LongAmp® Hot Start Taq DNA polymerase, Crimson LongAmp® Taq DNA polymerase, Crimson Taq DNA polymerase, OneTaq® DNA polymerase, OneTaq® Quick-Load® DNA polymerase, Hemo KlenTaq® DNA polymerase, REDTaq® DNA polymerase, Phusion® DNA polymerase, Phusion® High-Fidelity DNA polymerase, Platinum Pfx DNA polymerase, AccuPrime Pfx DNA polymerase, Klenow fragment, Pwo DNA polymerase, Pfu The 3' end of the oligonucleotide barcode is selected from the group consisting of a DNA polymerase, a T4 DNA polymerase, a T7 DNA polymerase, derivatives thereof, or any combination thereof. In some embodiments, extending the 3' end of the oligonucleotide barcode comprises extending the 3' end of the oligonucleotide barcode using a mesophilic DNA polymerase, a thermophilic DNA polymerase, a psychrophilic DNA polymerase, or any combination thereof. In some embodiments, extending the 3' end of the oligonucleotide barcode comprises extending the 3' end of the oligonucleotide barcode using a DNA polymerase lacking at least one of 5' to 3' exonuclease activity and 3' to 5' exonuclease activity, wherein the DNA polymerase may comprise a Klenow fragment. In some embodiments, extending the first plurality of oligonucleotide barcodes comprises extending the first plurality of oligonucleotide barcodes using a reverse transcriptase. In some embodiments, the reverse transcriptase is capable of terminal transferase activity.In some embodiments, the reverse transcriptase with strand displacement activity is PrimeScript reverse transcriptase, M-MuLV reverse transcriptase, SmartScribe reverse transcriptase, Maxima H Minus reverse transcriptase, and / or Superscript II reverse transcriptase. In some embodiments, the reverse transcriptase comprises a viral reverse transcriptase, which may be murine leukemia virus (MLV) reverse transcriptase or Moloney murine leukemia virus (MMLV) reverse transcriptase.

[0139] In some embodiments, each oligonucleotide barcode of the second plurality of oligonucleotide barcodes comprises a second molecular label, and at least 10 of the second plurality of oligonucleotide barcodes comprise a different second molecular label sequence, and each second molecular label may comprise at least 6 nucleotides, and further, the second molecular label sequence may be a random sequence. In some embodiments, the second plurality of oligonucleotide barcodes hybridize to the barcoded nucleic acid molecule by hybridization between the second molecular label and a sequence complementary to at least a portion of the nucleic acid target. In some embodiments, each oligonucleotide barcode of the second plurality of oligonucleotide barcodes comprises a second target binding region. In some embodiments, the first target binding region and / or the second target binding region comprises a poly(dA) region, a poly(dT) region, a random sequence, a gene-specific sequence, or any combination thereof. In some embodiments, the oligonucleotide barcodes of the second plurality of oligonucleotide barcodes hybridize to the barcoded nucleic acid molecule by hybridization between the second target binding region and a sequence complementary to at least a portion of the nucleic acid target. In some embodiments, at least ten of the second plurality of oligonucleotide barcodes comprise different second target binding regions, and at least two of the target binding regions may be capable of binding to the complement of different nucleic acid targets, and further, at least two of the target binding regions may be capable of hybridizing to different regions of the complement of the same nucleic acid target. In some embodiments, two or more oligonucleotide barcodes of the second plurality of oligonucleotide barcodes may hybridize to different regions of the complement of the same nucleic acid target to generate two or more extended barcoded nucleic acid molecules. In some embodiments, the two or more extended barcoded nucleic acid molecules may be generated by two or more oligonucleotide barcodes of the second plurality of oligonucleotide barcodes hybridizing to different regions of the complement of the same nucleic acid target.In some embodiments, the two or more extended barcoded nucleic acid molecules collectively comprise at least about 50% of the entire sequence of the nucleic acid target.

[0140] In some embodiments, the method includes denaturing a plurality of barcoded nucleic acid molecules. In some embodiments, the method includes denaturing a plurality of extended barcoded nucleic acid molecules. In some embodiments, the method includes determining the copy number of a nucleic acid target in a sample based on the number of first molecular labels having distinct sequences associated with the plurality of barcoded nucleic acid molecules or products thereof. In some embodiments, the method includes determining the copy number of a nucleic acid target in a sample based on the number of first molecular labels having distinct sequences, second molecular labels having distinct sequences, or a combination thereof associated with the plurality of extended barcoded nucleic acid molecules or products thereof. In some embodiments, determining the copy number of the nucleic acid target comprises determining the copy number of each of the plurality of nucleic acid targets in the sample based on the number of first molecular labels with distinct sequences associated with a barcoded nucleic acid molecule of the plurality of barcoded nucleic acid molecules, or a product thereof, that comprises the sequence of each of the plurality of nucleic acid targets; and / or the number of first molecular labels with distinct sequences, second molecular labels with distinct sequences, or a combination thereof, associated with an extended barcoded nucleic acid molecule of the plurality of extended barcoded nucleic acid molecules, that comprises the sequence of each of the plurality of nucleic acid targets. In some embodiments, the sequence of each of the plurality of nucleic acid targets comprises a subsequence of each of the plurality of nucleic acid targets. In some embodiments, the sequence of a nucleic acid target within the plurality of barcoded nucleic acid molecules comprises a subsequence of a nucleic acid target. In some embodiments, the nucleic acid target comprises mRNA. In some embodiments, the sample comprises a single cell, optionally an immune cell, and further optionally a B cell or a T cell. In some embodiments, the sample comprises a plurality of cells, a plurality of single cells, a tissue, a tumor sample, or any combination thereof. In some embodiments, the single cell comprises a circulating tumor cell.

[0141] In some embodiments, the first universal sequence of each oligonucleotide barcode of the first plurality of oligonucleotide barcodes is 5' to the first molecular label and the first target binding region; and / or the second universal sequence of each oligonucleotide barcode of the second plurality of oligonucleotide barcodes is 5' to the second molecular label and / or the second target binding region.

[0142] In some embodiments, the method includes amplifying a plurality of barcoded nucleic acid molecules using an amplification primer and a primer comprising a first universal sequence or a portion thereof, thereby generating a first plurality of single-labeled nucleic acid molecules comprising a sequence or a portion thereof of the nucleic acid target, and determining the copy number of the nucleic acid target in the sample includes determining the copy number of the nucleic acid target in the sample based on the number of first molecular labels having distinct sequences associated with the first plurality of single-labeled nucleic acid molecules or products thereof. In some embodiments, the method includes amplifying a plurality of extended barcoded nucleic acid molecules using an amplification primer and a primer comprising a second universal sequence or a portion thereof, thereby generating a second plurality of single-labeled nucleic acid molecules comprising a sequence or a portion thereof of the nucleic acid target, and determining the copy number of the nucleic acid target in the sample includes determining the copy number of the nucleic acid target in the sample based on the number of second molecular labels having distinct sequences associated with the second plurality of single-labeled nucleic acid molecules or products thereof. In some embodiments, the amplification primer comprises a fourth universal sequence. In some embodiments, the amplification primer is a target-specific primer. In some embodiments, the target-specific primers specifically hybridize to an immune receptor, a constant region of an immune receptor, a variable region of an immune receptor, a diversity region of an immune receptor, and / or a junction between a variable region and a diversity region of an immune receptor. In some embodiments, the immune receptor is a TCR and / or a BCR receptor, where the TCR may comprise a TCR alpha chain, a TCR beta chain, a TCR gamma chain, a TCR delta chain, or any combination thereof, and the BCR receptor comprises a BCR heavy chain and / or a BCR light chain.

[0143] In some embodiments, the method includes hybridizing random primers to a plurality of barcoded nucleic acid molecules and extending the random primers to generate a first plurality of extension products, where the random primers comprise a third universal sequence or its complement; and amplifying the first plurality of extension products using a primer capable of hybridizing to the third universal sequence or its complement and a primer capable of hybridizing to the first universal sequence or its complement, thereby generating a third plurality of single-labeled nucleic acid molecules or products thereof. In some embodiments, determining the copy number of the nucleic acid target in the sample includes determining the copy number of the nucleic acid target in the sample based on the number of first molecular labels having distinct sequences associated with the third plurality of single-labeled nucleic acid molecules or products thereof. In some embodiments, the method includes hybridizing a random primer to the plurality of extended barcoded nucleic acid molecules and extending the random primer to generate a second plurality of extension products, where the random primer comprises a third universal sequence or its complement; and amplifying the second plurality of extension products using a primer capable of hybridizing to the third universal sequence or its complement and a primer capable of hybridizing to the second universal sequence or its complement, thereby generating a fourth plurality of single-labeled nucleic acid molecules or products thereof. In some embodiments, determining the copy number of the nucleic acid target in the sample includes determining the copy number of the nucleic acid target in the sample based on the number of second molecular labels having distinct sequences associated with the fourth plurality of single-labeled nucleic acid molecules or products thereof.

[0144] In some embodiments, the first universal sequence, the second universal sequence, the third universal sequence, and / or the fourth universal sequence are the same. In some embodiments, the first universal sequence, the second universal sequence, the third universal sequence, and / or the fourth universal sequence are different. In some embodiments, the first universal sequence, the second universal sequence, the third universal sequence, and / or the fourth universal sequence comprise the binding site of the sequencing primer and / or the sequencing adaptor, their complementary sequence, and / or a portion thereof. In some embodiments, the sequencing adaptor comprises a P5 sequence, a P7 sequence, their complementary sequence, and / or a portion thereof. In some embodiments, the sequencing primer comprises a lead 1 sequencing primer, a lead 2 sequencing primer, their complementary sequence, and / or a portion thereof.

[0145] In some embodiments, the method includes obtaining sequence information of a plurality of barcoded nucleic acid molecules or products thereof. In some embodiments, obtaining sequence information includes attaching a sequencing adapter to a plurality of extended barcoded nucleic acid molecules or products thereof. In some embodiments, the method includes obtaining sequence information of a plurality of extended barcoded nucleic acid molecules or products thereof. In some embodiments, obtaining sequence information includes attaching a sequencing adapter to an extended barcoded nucleic acid molecule, a barcoded nucleic acid molecule, a product thereof, or any combination thereof. In some embodiments, the method includes obtaining sequence information of one or more of the first, second, third, and fourth plurality of single-labeled nucleic acid molecules or products thereof. In some embodiments, obtaining sequence information includes attaching a sequencing adapter to one or more of the first, second, third, and fourth plurality of single-labeled nucleic acid molecules or products thereof. In some embodiments, obtaining sequence information of one or more of the first, second, third, and fourth plurality of single-labeled nucleic acid molecules or products thereof comprises obtaining sequencing data comprising a plurality of sequencing reads of one or more of the first, second, third, and fourth plurality of single-labeled nucleic acid molecules or products thereof, each of the plurality of sequencing reads comprising (1) a cellular label sequence, (2) a molecular label sequence, and / or (3) a subsequence of the nucleic acid target.

[0146] In some embodiments, the method includes aligning each of a plurality of sequencing reads of the nucleic acid target with each unique cell label sequence indicative of a single cell of the sample to generate an aligned sequence of the nucleic acid target. In some embodiments, the aligned sequence of the nucleic acid target includes at least 50% of the cDNA sequence of the nucleic acid target, at least 70% of the cDNA sequence of the nucleic acid target, at least 90% of the cDNA sequence of the nucleic acid target, or the full length of the cDNA sequence of the nucleic acid target. In some embodiments, the nucleic acid target is an immune receptor, and the immune receptor may include a BCR light chain, a BCR heavy chain, a TCR alpha chain, a TCR beta chain, a TCR gamma chain, a TCR delta chain, or any combination thereof. In some embodiments, the aligned sequence of the nucleic acid target includes a complementarity determining region 1 (CDR1), a complementarity determining region 2 (CDR2), a complementarity determining region 3 (CDR3), a variable region, a full length of a variable region, or a combination thereof. In some embodiments, the aligned sequences of the nucleic acid targets include variable regions, diversity regions, junctions of variable regions, diversity regions and / or constant regions, or any combination thereof. In some embodiments, obtaining sequence information includes obtaining sequence information of the BCR light chain and BCR heavy chain of the single cell, and the sequence information of the BCR light chain and BCR heavy chain may include the sequence of the complementarity determining region 1 (CDR1), CDR2, CDR3, or any combination thereof, of the BCR light chain and / or BCR heavy chain. In some embodiments, the method includes pairing the BCR light chain and BCR heavy chain of the single cell based on the obtained sequence information. In some embodiments, the sample includes a plurality of single cells, and the method includes pairing the BCR light chain and BCR heavy chain of at least 50% of the single cells based on the obtained sequence information. In some embodiments, obtaining sequence information comprises obtaining sequence information of the TCR alpha chain and the TCR beta chain of the single cell, and the sequence information of the TCR alpha chain and the TCR beta chain may comprise the sequence of complementarity determining region 1 (CDR1), CDR2, CDR3, or any combination thereof, of the TCR alpha chain and / or the TCR beta chain. In some embodiments, the method comprises pairing the TCR alpha chain and the TCR beta chain of the single cell based on the obtained sequence information.In some embodiments, the sample comprises a plurality of single cells, and the method comprises pairing the TCR alpha chain and the TCR beta chain of at least 50% of the single cells based on the sequence information obtained. In some embodiments, obtaining sequence information comprises obtaining sequence information of the TCR gamma chain and the TCR delta chain of the single cells. In some embodiments, the sequence information of the TCR gamma chain and the TCR delta chain comprises the sequence of the complementarity determining region 1 (CDR1), CDR2, CDR3, or any combination thereof, of the TCR gamma chain and / or the TCR delta chain. In some embodiments, the method comprises pairing the TCR gamma chain and the TCR delta chain of the single cells based on the sequence information obtained. In some embodiments, the sample comprises a plurality of single cells, and the method comprises pairing the TCR gamma chain and the TCR delta chain of at least 50% of the single cells based on the sequence information obtained.

[0147] In some embodiments, the complement of the molecular label comprises a reverse complement sequence of the molecular label or a complementary sequence of the molecular label. In some embodiments, the plurality of barcoded nucleic acid molecules comprises barcoded deoxyribonucleic acid (DNA) molecules, barcoded ribonucleic acid (RNA) molecules, or a combination thereof. In some embodiments, the nucleic acid target comprises a nucleic acid molecule, which may comprise ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNA comprising a poly(A) tail, or any combination thereof, and further, the mRNA may encode an immune receptor. In some embodiments, the nucleic acid target comprises a cellular component binding reagent, and / or the nucleic acid molecule is associated with a cellular component binding reagent, and the method may further comprise dissociating the nucleic acid molecule and the cellular component binding reagent. In some embodiments, at least 10 of the first and / or second plurality of oligonucleotide barcodes comprise different molecular label sequences. In some embodiments, each molecular label of the first and / or second plurality of oligonucleotide barcodes comprises at least 6 nucleotides. In some embodiments, the first and / or second plurality of oligonucleotide barcodes are associated with a solid support. In some embodiments, the first and / or second plurality of oligonucleotide barcodes associated with the same solid support each comprise the same sample label. In some embodiments, each sample label of the first and / or second plurality of oligonucleotide barcodes comprises at least 6 nucleotides. In some embodiments, the first and / or second plurality of oligonucleotide barcodes each comprise a cell label. In some embodiments, each cell label of the first and / or second plurality of oligonucleotide barcodes comprises at least 6 nucleotides. In some embodiments, the oligonucleotide barcodes of the first and / or second plurality of oligonucleotide barcodes associated with the same solid support comprise the same cell label. In some embodiments, the oligonucleotide barcodes of the first and / or second plurality of oligonucleotide barcodes associated with different solid supports comprise different cell labels.In some embodiments, the method includes extending the oligonucleotide barcode in the presence of one or more of ethylene glycol, polyethylene glycol, 1,2-propanediol, dimethylsulfoxide (DMSO), glycerol, formamide, 7-deaza-GTP, acetamide, tetramethylammonium chloride salts, betaine, or any combination thereof.

[0148] In some embodiments, the solid support comprises a synthetic particle, a planar surface, or a combination thereof. In some embodiments, the sample comprises a single cell, and the method comprises associating a synthetic particle comprising a first and a second plurality of oligonucleotide barcodes with the single cell in the sample. In some embodiments, the method comprises lysing the single cell after associating the synthetic particle with the single cell, where lysing the single cell may comprise heating the sample, contacting the sample with a detergent, changing the pH of the sample, or any combination thereof. In some embodiments, the synthetic particle and the single cell are in the same compartment, which may be a well or a droplet. In some embodiments, at least one oligonucleotide barcode of the first and / or second plurality of oligonucleotide barcodes is immobilized or partially immobilized on the synthetic particle, or at least one oligonucleotide barcode of the first and / or second plurality of oligonucleotide barcodes is encapsulated or partially encapsulated within the synthetic particle. In some embodiments, the synthetic particle is disintegrable, and may be a disintegrable hydrogel particle. In some embodiments, the synthetic particles comprise beads, which may be sepharose beads, streptavidin beads, agarose beads, magnetic beads, conjugated beads, protein A conjugated beads, protein G conjugated beads, protein A / G conjugated beads, protein L conjugated beads, oligo(dT) conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, or any combination thereof. In some embodiments, the synthetic particles comprise a material selected from the group consisting of polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic, ceramic, plastic, glass, methylstyrene, acrylic polymer, titanium, latex, sepharose, cellulose, nylon, silicone, and any combination thereof.In some embodiments, each oligonucleotide barcode of the first and / or second plurality of oligonucleotide barcodes comprises a linker functional group. In some embodiments, the synthetic particle comprises a solid support functional group. In some embodiments, the support functional group and the linker functional group are associated with each other, and the linker functional group and the support functional group may each be selected from the group consisting of C6, biotin, streptavidin, primary amines, aldehydes, ketones, and any combination thereof.

[0149] In some embodiments, a solid support is provided. Disclosed herein in some embodiments is a solid support associated with one or both of a first and second plurality of oligonucleotide barcodes.

[0150] The methods and systems described herein can be used in connection with methods and systems that use antibodies associated with (e.g., bound to or conjugated with) oligonucleotides (also referred to herein as AbOs or AbOligos). Some embodiments using AbOs to determine protein expression profiles in single cells and trace sample origin are described in US Patent Application Publication Nos. 2018 / 0088112 and 2018 / 0346970, the contents of each of which are incorporated herein by reference in their entireties. In some embodiments, the methods disclosed herein allow for T and B cell V(D)J profiling, 3' targeting, 5' targeting, 3' whole transcriptome amplification (WTA), 5' WTA, protein expression profiling using AbOs, and / or sample multiplexing in a single experiment. Methods for determining the sequence of a nucleic acid target (e.g., the V(D)J region of an immune receptor) using 5' barcoding and / or 3' barcoding are described in U.S. Patent Application Publication No. 2020 / 0109437, the contents of which are incorporated herein by reference in their entirety. Systems, methods, compositions, and kits for molecular barcoding at the 5' end of a nucleic acid target are described, for example, in U.S. Patent Application Publication No. 2019 / 0338278, the contents of which are incorporated herein by reference in their entirety.The systems, methods, compositions, and kits for internal-based gene expression profiling provided herein can, in some embodiments, be used in conjunction with methods for obtaining full-length V(D)J information using a combination of 5' barcoding and random priming approaches (e.g., by Illumina sequencing on a Rhapsody system) as described in U.S. patent application Ser. No. 17 / 091,639, entitled "USING RANDOM PRIMING TO OBTAIN FULL-LENGTH V(D)J INFORMATION FOR IMMUNE REPERTOIRE SEQUENCING," filed Nov. 6, 2020, the contents of which are incorporated herein by reference in their entirety. The systems, methods, compositions, and kits for internal-based gene expression profiling provided herein can, in some embodiments, be used with random priming and extension (RPE)-based whole transcriptome analysis methods and compositions described in U.S. Patent Application No. 16 / 677,012, the contents of which are incorporated herein by reference in their entirety. The systems, methods, compositions, and kits for internal-based gene expression profiling provided herein can, in some embodiments, be used with blocker oligonucleotides described in U.S. Patent Application Publication No. 20210238661, the contents of which are incorporated herein by reference in their entirety. Some embodiments of the compositions and methods disclosed herein include a first plurality of oligonucleotide barcodes and a second plurality of oligonucleotide barcodes, which are described in U.S. Patent Application Publication No. 20210371909, the contents of which are incorporated herein by reference in their entirety.The systems, methods, compositions, and kits provided herein can, in some embodiments, be used with template switch oligonucleotides containing blocking sequences as described in PCT Patent Application No. PCT / US22 / 75577, entitled "TEMPLATE SWITCH OLIGONUCLEOTIDE (TSO) FOR MRNA 5' ANALYSIS," filed August 29, 2022, the contents of which are incorporated herein by reference in their entirety, which in some embodiments can reduce the production of undesired extension products during library preparation.

[0151] In some embodiments, the extension product and / or amplification product disclosed herein can be used for sequencing.Any suitable sequencing method known in the art can be used, preferably high-throughput approach.For example, circular array sequencing can also be used, using platforms such as Roche454, Illumina Solexa, ABI-SOLiD, ION Torrent, Complete Genomics, Pacific Bioscience, Helicos or Polonator platform.Sequencing can include MiSeq sequencing and / or HiSeq sequencing.

[0152] The present disclosure includes systems, methods, compositions, and kits for attaching barcodes (e.g., stochastic barcodes) having molecular labels (or molecular indexes) to the 5' end of barcoded or labeled nucleic acid targets (e.g., deoxyribonucleic acid molecules and ribonucleic acid molecules). The 5'-based and internal-based transcript counting methods disclosed herein can complement or complement, for example, 3'-based transcript counting methods (e.g., Rhapsody™ Assay (Becton, Dickinson and Company, Franklin Lakes, NJ), Chromium™ Single Cell 3' Solution (10X Genomics, San Francisco, CA)). Barcoded nucleic acid targets can be used for sequence identification, transcript counting, alternative splicing analysis, mutation screening, and / or full-length sequencing in a high-throughput manner. Transcript counting at the 5' end (5' to the labeled target nucleic acid target) can reveal alternative splicing isoforms and variants at or near the 5' end of the nucleic acid molecule, including but not limited to splice variants, single nucleotide polymorphisms (SNPs), insertions, deletions, substitutions. In some embodiments, the method can involve intramolecular hybridization. Methods for determining the sequence of a nucleic acid target (e.g., the V(D)J region of an immune receptor) using 5' barcoding and / or 3' barcoding are described in US2020 / 0109437, the contents of which are incorporated herein by reference in their entirety. Systems, methods, compositions, and kits for molecular barcoding at the 5' end of a nucleic acid target are described in US Patent Application Publication No. 2019 / 0338278, the contents of which are incorporated herein by reference in their entirety.

[0153] The disclosed method can be used to identify the VDJ regions of BCR, TCR, and antibodies. VDJ recombination, also known as somatic recombination, is a mechanism of genetic recombination in the early stages of immune system immunoglobulin (Ig) (e.g., BCR) and TCR production. VDJ recombination allows variable (V), diversity (D) and joining (J) gene segments to be combined in a near-random manner. The randomness in selecting various genes allows for a variety of coding proteins that match antigens from bacteria, viruses, parasites, dysfunctional cells such as tumor cells, and pollen.

[0154] The VDJ region may contain a large locus of 3 Mb that contains variable (V), diversity (D) and joining (J) genes. These are the segments that may be involved in VDJ recombination. There may also be constant genes that may not undergo VDJ recombination. The first event in the VDJ recombination of this locus may be the rearrangement of one of the D genes to one of the J genes. After this, one of the V genes can be added to this DJ rearrangement to form a functional VDJ rearranged gene that later encodes the variable segment of the heavy chain protein. Both of these steps may be catalyzed by recombinase enzymes that may delete the intervening DNA.

[0155] This recombination process occurs stepwise in precursor B cells, resulting in the diversity required for the antibody repertoire. Each B cell can produce only one antibody (e.g., BCR). This specificity can be achieved by allelic exclusion, such that the functional rearrangement of one allele transmits a signal to prevent further recombination of the second allele. In some embodiments, the sample comprises immune cells, which may include, for example, T cells, B cells, lymphoid stem cells, myeloid progenitor cells, lymphocytes, granulocytes, B cell precursors, T cell precursors, natural killer cells, Tc cells, Th cells, plasma cells, memory cells, neutrophils, eosinophils, basophils, mast cells, monocytes, dendritic cells, and / or macrophages, or any combination thereof.

[0156] T cells may be derived from a single T cell or a T cell clone, which may refer to T cells with the same TCR. T cells may be part of a T cell line, which may include a mixed population of T cell clones and T cells with different TCRs, but all of which may recognize the same target (e.g., antigen, tumor, virus). T cells can be obtained from several sources, including peripheral blood mononuclear cells, bone marrow, lymph node tissue, splenic tissue, and tumors. T cells can be obtained from a unit of blood taken from a subject, such as using Ficoll separation. Cells derived from an individual's circulating blood can be obtained by apheresis or leukapheresis. Apheresis products may include T cells, monocytes, granulocytes, lymphocytes including B cells, other nucleated white blood cells, red blood cells, and platelets. The cells can be washed and resuspended in medium to isolate the cells of interest.

[0157] T cells can be isolated from peripheral blood lymphocytes by lysing red blood cells and depleting monocytes, for example, by centrifugation through a PERCOLL™ gradient. Specific subpopulations of T cells, for example, CD28+, CD4, CDC, CD45RA+, and CD45RO+ T cells, can be further isolated by positive or negative selection techniques. For example, T cells can be isolated by incubation with anti-CD3 / anti-CD28 (i.e., 3×28) conjugated beads, for example, DYNABEADS® M-450 CD3 / CD28 T, or XCYTE DYNABEADS™, for a time sufficient for positive selection of the desired T cells. Immune cells (e.g., T cells and B cells) can be antigen-specific (e.g., specific for tumors). In some embodiments, the cell may be an antigen presenting cell (APC), such as a B cell, an activated B cell from a lymph node, a lymphoblastoid cell, a resting B cell, or a neoplastic B cell, for example, from a lymphoma. APC may refer to a B cell or a follicular dendritic cell that expresses at least one of the BCRC proteins on its surface.

[0158] The disclosed methods can be used to track the molecular phenotype of single T cells. Different subtypes of T cells can be distinguished by the expression of different molecular markers. T cells express unique TCRs from a diverse repertoire of TCRs. In most T cells, the TCR may be composed of a heterodimer of α and β chains, and each functional chain may be the product of somatic DNA recombination events during T cell development, allowing the expression of over one million different TCRs in a single individual. The TCRs can be used to define the identity of individual T cells, allowing lineage tracing of T cell clonal expansion during an immune response. The disclosed immunological methods can be used in a variety of ways, including but not limited to, identifying unique TCR α and TCR β chain pairings in single T cells, quantifying TCR and marker expression at the single cell level, identifying TCR diversity in an individual, characterizing the TCR repertoire expressed in different T cell populations, determining the functionality of TCR alpha and beta chain alleles, and identifying clonal expansion of T cells during an immune response.

[0159] T cell receptor chain pairing TCR is a recognition molecule present on the surface of T lymphocytes. T cell receptors found on the surface of T cells can be composed of two glycoprotein subunits called alpha and beta chains. Both chains contain a molecular weight of about 40 kDa and can have variable and constant domains. Genes encoding alpha and beta chains can be organized in libraries of V, D and J regions where genes are formed by gene rearrangement. TCR can recognize antigens presented by antigen presenting cells as part of a complex with specific self molecules encoded by histocompatibility genes. The most prevalent histocompatibility genes are known as major histocompatibility complexes (MHC). Thus, the complex recognized by the T cell receptor consists of MHC / peptide ligand.

[0160] In some embodiments, the disclosed methods, devices, and systems can be used for sequencing and pairing of TCRs. The disclosed methods, devices, and systems can be used to sequence T cell receptor alpha and beta chains, pair alpha and beta chains, and / or determine functional copies of T cell receptor alpha chains. A single cell can be contained in a single compartment (e.g., well) that contains a single solid support (e.g., bead). The cell can be lysed. The bead can contain stochastic labels that can bind to specific positions in the TCR alpha and / or beta chains. The TCR alpha and beta molecules associated with the solid support can be subjected to the disclosed molecular biology methods, including reverse transcription, amplification, and sequencing. TCR alpha and beta chains that contain the same cell label are considered to be derived from the same single cell, thereby allowing the TCR alpha and beta chains to be paired.

[0161] Heavy- and light-chain pairing in the antibody repertoire The disclosed methods, devices and systems can be used for pairing of receptors of BCR and heavy and light chains of antibodies. The disclosed methods allow the repertoire of immune receptors and antibodies in an individual organism or cell population to be determined. The disclosed methods can help determine the pairs of polypeptide chains that make up immune receptors. B cells and T cells each express immune receptors, B cells express immunoglobulins and BCRs, and T cells express TCRs. Both types of immune receptors can include two polypeptide chains. Immunoglobulins can include variable heavy (VH) and variable light (VL) chains. There are two types of TCRs, one consisting of alpha and beta chains, and one consisting of delta and gamma chains. Polypeptides in immune receptors can include constant and variable regions. The variable regions can result from recombination and end-joining rearrangements of gene fragments in the chromosomes of B or T cells. In B cells, further diversification of the variable regions can occur by somatic hypermutation.

[0162] The immune system has a large repertoire of receptors, and any given pair of receptors expressed by a lymphocyte may be encoded by a separate, unique pair of transcripts. Knowledge of the sequences of immune receptor chain pairs expressed in a single cell can be used to ascertain the immune repertoire of a given individual or cell population. In some embodiments, the disclosed methods, devices, and systems can be used for antibody sequencing and pairing. The disclosed methods, devices, and systems can be used for antibody heavy and light chain sequencing (e.g., in B cells) and / or heavy and light chain pairing. A single cell can be contained in a single compartment (e.g., well) that includes a single solid support (e.g., bead). The cell can be lysed. The beads may contain stochastic labels that can bind to specific locations within the heavy and / or light chains of the antibody (e.g., in B cells). The heavy and light chain molecules associated with the solid support can be subjected to the molecular biology methods of the present disclosure, including reverse transcription, amplification, and sequencing. The heavy and light chains of the antibody that contain the same cell label are considered to be derived from the same single cell, thereby allowing the heavy and light chains of the antibody to be paired.

[0163] The methods disclosed herein can allow for 3'-, internal-, and / or 5'-based sequencing. The methods can allow for flexibility in sequencing. In some embodiments, the methods can allow for profiling of immune repertoires of both T and B cells in the Rhapsody™ system for samples, e.g., mouse and human samples. In some embodiments, 3', internal, and / or 5' expression profiling of V(D)J can be performed. In some embodiments, both phenotypic markers and V(D)J sequences of T and B cells in a single cell platform can be investigated. In some embodiments, 3', internal, and 5' information of those transcripts can be captured in a single experiment. The methods disclosed herein can allow for V(D)J detection (e.g., hypermutation) of both T and B cells.

[0164] The methods and systems described herein can be used with respect to methods and systems that use antibodies associated with (e.g., bound to or conjugated to) oligonucleotides (also referred to herein as AbOs or AbOligos). Embodiments using AbOs to determine protein expression profiles in single cells and trace sample origin are described in U.S. Patent Application No. 15 / 715,028, published as U.S. Patent Application Publication No. 2018 / 0088112, and U.S. Patent Application No. 15 / 937,713, the contents of each of which are incorporated herein by reference in their entireties. In some embodiments, the methods disclosed herein allow for T and B cell V(D)J profiling, 3' targeting, 5' targeting, 3' whole transcriptome amplification (WTA), 5' WTA, protein expression profiling with AbOs, and / or sample multiplexing in a single experiment.

[0165] In some embodiments, the step of extending the random primer is performed at a substantially constant temperature. In some embodiments, the step of extending the random primer is performed at a constant temperature. In some embodiments, the step of extending the random primer is started at a first extension temperature. In some embodiments, the step of extending the random primer is performed at one or more temperatures different from the first extension temperature (e.g., a second extension temperature and / or a third extension temperature). The second extension temperature and / or the third extension temperature may be higher or lower than the first extension temperature. In some embodiments, the first extension temperature and / or the second extension temperature is about 30°C, 31°C, 32°C, 33°C, 34°C, 35°C, 36°C, 37°C, 38°C, 39°C, 40°C, 41°C, 42°C, 43°C, 44°C, 45°C, 46°C, 47°C, 48°C, 49°C, 50°C, 51°C, 52°C, 53°C, 54°C, 55°C, 56°C, 57°C, 58°C, 59°C, 60°C, 61°C, 62°C, 63°C, 64°C, 65°C, 66°C, 67°C, 68°C, 69°C, 70°C, 71°C, 72°C, 73°C, 74°C, 75°C, 76°C, 77°C, 78°C, 79°C, 80°C, or a number or range between any two of these values. In some embodiments, the first extension temperature is about 37° C. In some embodiments, the second extension temperature is about 55° C. In some embodiments, the second extension temperature is about 45° C.

[0166] The number of cycles of random priming and extension may vary in different implementations. In some embodiments, the number of cycles of random priming and extension is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 60, 70, 80, 90, 100, or more. It may include a number or range of random priming and extension cycles between any two of these values, it may include a number or range of random priming and extension cycles that is approximately these values ​​or between any two of these values, it may include a number or range of random priming and extension cycles that is at least these values ​​or between any two of these values, or it may include a number or range of random priming and extension cycles that is at most these values ​​or between any two of these values.

[0167] The random primer may comprise a random sequence of nucleotides. The random sequence of nucleotides may be about 4 to about 30 nucleotides in length. In some embodiments, the random sequence of nucleotides is 6 or 9 nucleotides in length. The random sequence of nucleotides may have different lengths in different implementations. In some embodiments, the random sequence of nucleotides in the random primer is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, or a number or range of nucleotides between any two of these values ​​in length, is about, is at least, or is up to the number or range of nucleotides between any two of these values ​​in length. The random primers may have different concentrations during the random priming step in different implementations. In some embodiments, the random primers are at a concentration of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 128, or any two of these values, or a number or range of uM during random priming, or at about, or at least, or at most, or any two of these values, or a number or range of uM.

[0168] Methods for producing oligonucleotide barcodes and barcoded particles are described, for example, in U.S. Patent Application Publication No. 2015 / 0299784, International Publication No. WO2015 / 031691, Fu et al, PNAS USA 2011 May 31;108(22):9026-31, and U.S. Patent Application No. 17 / 336,055, filed June 1, 2021, entitled “OLIGONUCLEOTIDES AND BEADS FOR 5 PRIME GENE EXPRESSION ASSAY,” the contents of which are incorporated herein in their entireties.

[0169] While various aspects and embodiments are disclosed herein, other aspects and embodiments will be apparent to those of ordinary skill in the art. The various aspects and embodiments disclosed herein are intended to be illustrative and not limiting, with the true scope and spirit being indicated by the following claims. Those skilled in the art will recognize that for this and other processes and methods disclosed herein, the functions performed in the processes and methods may be realized in differing orders. Moreover, the outlined steps and operations are provided only as examples, and some of the steps and operations may be arbitrarily combined into fewer steps and operations or expanded into additional steps and operations without diminishing the essential elements of the embodiments of the present disclosure.

[0170] In connection with the use of substantially any plural and / or singular terms herein, those skilled in the art can convert from plural to singular and / or from singular to plural, where appropriate in the context and / or application. Various singular / plural permutations may be expressly set forth herein for the sake of clarity.

[0171] In general, it will be understood by those skilled in the art that the terms used herein, and particularly in the appended claims (e.g., the body of the appended claims), are generally intended to be "open" terms (e.g., the term "including" should be interpreted as "including but not limited to," the term "having" should be interpreted as "having at least," the term "includes" should be interpreted as "includes but is not limited to," etc.). Furthermore, it will be understood by those skilled in the art that where a particular number of introduced claim recitations are intended, such intent will be expressly set forth in the claim, and that in the absence of such recitation, no such intent exists. For example, as an aid to understanding, the following appended claims may include the use of the introductory phrases "at least one" and "one or more" to introduce the claim recitations. However, the use of such phrases should not be construed as meaning that the introduction of a claim recitation with the indefinite article "a" or "an" limits any particular claim that includes such an introduced claim recitation to embodiments containing only one such recitation, even if the same claim also includes the introductory phrase "one or more" or "at least one" and an indefinite article such as "a" or "an" (e.g., "a" and / or "an" should be construed to mean "at least one" or "one or more"), nor should the use of definite articles used to introduce claim recitations.Moreover, even if a particular number of introduced claim recitations is explicitly recited, one of ordinary skill in the art will recognize that such recitation should be interpreted to mean at least the recited number (e.g., an unmodified recitation such as "two recitations" without other modifiers means at least two recitations, or more than two recitations). Furthermore, when a convention similar to "such as at least one of A, B, and C" is used, such a configuration is generally intended to mean as one of ordinary skill in the art would understand the convention (e.g., "a system having at least one of A, B, and C" would include, but is not limited to, systems having A alone, B alone, C alone, both A and B, both A and C, both B and C, and / or all of A, B, and C, etc.). When a convention similar to "such as at least one of A, B, or C" is used, such construction is generally intended to be in the sense that one of skill in the art would understand that convention (e.g., "a system having at least one of A, B, or C" would include, but is not limited to, systems having A alone, B alone, C alone, both A and B, both A and C, both B and C, and / or all of A, B, and C, etc.). Furthermore, it will be understood by those of skill in the art that virtually any disjunctive word and / or phrase expressing two or more alternative terms, regardless of the specification, claims, or drawings, should be understood to contemplate the possibility of including one of the terms, either of the terms, or both of the terms. For example, the phrase "A or B" is understood to include the possibilities of "A" or "B" or "A and B."

[0172] Furthermore, when features or aspects of the disclosure are described in terms of a Markush group, it will be recognized by those of skill in the art that the disclosure is also described in terms of any individual members or subgroups of members of the Markush group.

[0173] As will be understood by one of ordinary skill in the art, for all purposes, e.g., with respect to the provision of the specification, all ranges disclosed herein encompass all possible subranges and combinations of subranges thereof. Any recited range is readily identifiable as being fully descriptive and that the range can be divided into at least 2, 3, 4, 5, 10, etc. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third, upper third, etc. Similarly, as will be understood by one of ordinary skill in the art, all expressions such as "up to," "at least," etc. refer to ranges that are inclusive of the recited numbers and that can be subsequently broken down into subranges as discussed above. Finally, as will be understood by one of ordinary skill in the art, ranges include each individual member. Thus, for example, a group having 1-3 cells refers to a group having 1, 2, or 3 cells. Similarly, a group having 1-5 cells refers to a group having 1, 2, 3, 4, or 5 cells, and so forth.

[0174] From the foregoing, it will be appreciated that various embodiments of the present disclosure have been described herein for purposes of illustration, and that various modifications may be made without departing from the scope and spirit of the present disclosure. Accordingly, the various embodiments disclosed herein are not intended to be limiting, with the true scope and spirit being indicated by the following claims.

Claims

1. 1. A method for labeling nucleic acid targets in a sample, comprising: contacting copies of the nucleic acid target with a first plurality of oligonucleotide barcodes, each oligonucleotide barcode of the first plurality of oligonucleotide barcodes comprising a first universal sequence, a first molecular label, and a first target binding region capable of hybridizing to the nucleic acid target; extending the first plurality of oligonucleotide barcodes hybridized to copies of the nucleic acid target to generate a plurality of barcoded nucleic acid molecules, each comprising a first universal sequence, a first molecular label, and a sequence complementary to at least a portion of the nucleic acid target; contacting the barcoded nucleic acid molecules with a second plurality of oligonucleotide barcodes for hybridization; each oligonucleotide barcode of the second plurality of oligonucleotide barcodes comprises a second universal sequence, a cleavage domain, and a blocking group; a blocking group capable of preventing extension of the oligonucleotide barcode, and a cleavage domain located 5' to the blocking group, such that when the cleavage domain hybridizes to the barcoded nucleic acid molecule, the oligonucleotide barcode can be cleaved by a cleavage enzyme at a point within or adjacent to the cleavage domain; contacting a second plurality of oligonucleotide barcodes hybridized to the barcoded nucleic acid molecules with a cleavage enzyme, thereby removing blocking groups from said oligonucleotide barcodes; extending 3' ends of oligonucleotide barcodes of the second plurality of oligonucleotide barcodes hybridized to the barcoded nucleic acid molecules to generate a plurality of extended barcoded nucleic acid molecules; The above method, comprising:

2. A method according to claim 1, a. each extended barcoded nucleic acid molecule of the plurality of extended barcoded nucleic acid molecules comprises the sequence of at least a portion of a nucleic acid target; and / or b. the cleavage enzyme is an RNase H enzyme and / or an RNase H2 enzyme, optionally the RNase H2 enzyme may be a Pyrococcus abyssi RNase H2 enzyme; and / or c. the cleavage enzyme is a hot-start cleavage enzyme that is thermostable and has reduced activity at low temperatures, and optionally the hot-start cleavage enzyme may be Pyrococcus abysilibonuclease H2 containing: (a) a G12A amino acid substitution; (b) a P13T amino acid substitution; (c) a G169A amino acid substitution; or (d) a combination thereof; and / or d. the cleavage enzyme is chemically modified, and / or e. The method according to claim 1, wherein the cleavage enzyme is a chemically modified hot-start cleavage enzyme that is thermostable and has reduced activity at low temperatures, and optionally, the cleavage enzyme is reversibly inactivated at low temperatures by interaction with an antibody.

3. A method according to claim 1, a. the cleavage domain comprises one or more ribonucleotides that are cleavable by an RNase H enzyme; and / or b. the cleavage domain comprises one or more of the following moieties: a DNA residue, an abasic residue, a modified nucleoside, or a modified phosphate internucleotide linkage; and / or c. the cleavage domain comprises at least one RNA base, and / or d. The method of any of the preceding claims, wherein the cleavage domain comprises one or more 2'-modified nucleosides, and optionally, one or more modified nucleosides can be 2'-fluoronucleosides.

4. The method of claim 1, a. a blocking group is attached to the 3' terminal nucleotide of the oligonucleotide barcode, and / or b. the blocking group is at or near the 3' end of the oligonucleotide barcode, and / or c. the blocking group is a 2',3'-dideoxynucleotide, a ribonucleotide residue, a 2',3'-SH nucleotide, or a 2'-O-PO trinucleotide; and / or d. the blocking group comprises a non-nucleotide modification, and / or e. The method according to any one of claims 1 to 5, wherein the blocking group further comprises a naphthyl-azo compound, a spacer, and / or biotin.

5. The method of claim 1, a. extending 3'-ends of oligonucleotide barcodes of the second plurality of oligonucleotide barcodes comprises extending 3'-ends of oligonucleotide barcodes of the second plurality of oligonucleotide barcodes using a DNA polymerase having strand displacement activity; Optionally, extending a 3′ end of an oligonucleotide barcode of the second plurality of oligonucleotide barcodes using a DNA polymerase with strand displacement activity may be capable of generating an extended barcoded nucleic acid molecule comprising a complement of the first molecular label and a complement of the first universal sequence; and / or b. extending the 3' ends of the oligonucleotide barcodes of the second plurality of oligonucleotide barcodes comprises extending the 3' ends of the oligonucleotide barcodes of the second plurality of oligonucleotide barcodes using a DNA polymerase that does not have strand displacement activity; and / or c. The polymerase is selected from the group consisting of Phi29 DNA polymerase, E. coli DNA polymerase I, Bsu DNA polymerase, Bst DNA polymerase, Taq DNA polymerase, VENT™ DNA polymerase, DEEPVENT™ DNA polymerase, LongAmp® Taq DNA polymerase, LongAmp® Hot Start Taq DNA polymerase, Crimson LongAmp® Taq DNA polymerase, Crimson Taq DNA polymerase, OneTaq® DNA polymerase, OneTaq® Quick-Load® DNA polymerase, Hemo and / or selected from the group consisting of KlenTaq® DNA polymerase, REDTaq® DNA polymerase, Phusion® DNA polymerase, Phusion® High-Fidelity DNA polymerase, Platinum Pfx DNA polymerase, AccuPrime Pfx DNA polymerase, Klenow fragment, Pwo DNA polymerase, Pfu DNA polymerase, T4 DNA polymerase, T7 DNA polymerase, derivatives thereof, or any combination thereof; d. extending the 3' end of the oligonucleotide barcode comprises extending the 3' end of the oligonucleotide barcode using a mesophilic DNA polymerase, a thermophilic DNA polymerase, a psychrophilic DNA polymerase, or any combination thereof; and / or e. extending the 3' end of the oligonucleotide barcode comprises extending the 3' end of the oligonucleotide barcode using a DNA polymerase that lacks at least one of 5' to 3' exonuclease activity and 3' to 5' exonuclease activity, optionally wherein the DNA polymerase comprises Klenow fragment; and / or f. extending the first plurality of oligonucleotide barcodes comprises extending the first plurality of oligonucleotide barcodes using a reverse transcriptase; Optionally, i. the reverse transcriptase may be capable of terminal transferase activity, and / or ii. The reverse transcriptase with strand displacement activity may be PrimeScript reverse transcriptase, M-MuLV reverse transcriptase, SmartScribe reverse transcriptase, Maxima H Minus reverse transcriptase, and / or Superscript II reverse transcriptase; and / or iii. The method of any of the preceding claims, wherein the reverse transcriptase comprises a viral reverse transcriptase, and optionally, the viral reverse transcriptase may be murine leukemia virus (MLV) reverse transcriptase or Moloney murine leukemia virus (MMLV) reverse transcriptase.

6. The method of claim 1, a. each oligonucleotide barcode of the second plurality of oligonucleotide barcodes comprises a second molecular label, and at least 10 of the second plurality of oligonucleotide barcodes comprise different second molecular label sequences, optionally, each second molecular label comprises at least 6 nucleotides, and further optionally, the second molecular label sequences may be random sequences; and / or b. a second plurality of oligonucleotide barcodes hybridize to the barcoded nucleic acid molecules by hybridization between the second molecular label and a sequence complementary to at least a portion of the nucleic acid target; and / or c. each oligonucleotide barcode of the second plurality of oligonucleotide barcodes comprises a second target binding region; and / or d. the first target binding region and / or the second target binding region comprises a poly(dA) tract, a poly(dT) tract, a random sequence, a gene-specific sequence, or any combination thereof; and / or e. each oligonucleotide barcode of the second plurality of oligonucleotide barcodes hybridizes to a barcoded nucleic acid molecule by hybridization between the second target binding region and a sequence complementary to at least a portion of the nucleic acid target; and / or f. at least 10 of the second plurality of oligonucleotide barcodes comprise different second target binding regions, and optionally, at least two of the target binding regions may be capable of binding to the complement of different nucleic acid targets, and further optionally, at least two of the target binding regions may be capable of hybridizing to different regions of the complement of the same nucleic acid target; and / or g. two or more oligonucleotide barcodes of the second plurality of oligonucleotide barcodes may be capable of hybridizing to different regions of the complement of the same nucleic acid target to generate two or more elongated barcoded nucleic acid molecules; Optionally, i. the two or more elongated barcoded nucleic acid molecules may be generated by two or more oligonucleotide barcodes of a second plurality of oligonucleotide barcodes that hybridize to different regions of the complement of the same nucleic acid target; and / or ii. The method of any of the above, wherein the two or more extended barcoded nucleic acid molecules collectively comprise at least about 50% of the entire sequence of the nucleic acid target.

7. The method of claim 1, a. the method comprises denaturing a plurality of barcoded nucleic acid molecules; and / or b. the method comprises denaturing the plurality of elongated barcoded nucleic acid molecules; and / or c. the method further comprises determining the copy number of the nucleic acid target in the sample based on the number of first molecular labels having distinct sequences associated with the plurality of barcoded nucleic acid molecules or products thereof; Optionally, the step of determining the copy number of the nucleic acid target comprises: i. the number of first molecular labels having distinct sequences associated with barcoded nucleic acid molecules, or products thereof, among the plurality of barcoded nucleic acid molecules that comprise the sequences of each of the plurality of nucleic acid targets; and / or ii. the number of first molecular labels having distinct sequences, second molecular labels having distinct sequences, or combinations thereof, associated with extended barcoded nucleic acid molecules of the plurality of extended barcoded nucleic acid molecules that comprise the sequences of each of the plurality of nucleic acid targets; and / or determining the copy number of each of the plurality of nucleic acid targets in the sample based on d. the method further comprises determining the copy number of the nucleic acid target in the sample based on the number of first molecular labels having distinct sequences, second molecular labels having distinct sequences, or a combination thereof, associated with a plurality of the extended barcoded nucleic acid molecules or products thereof; Optionally, the step of determining the copy number of the nucleic acid target comprises: i. the number of first molecular labels having distinct sequences associated with barcoded nucleic acid molecules, or products thereof, among the plurality of barcoded nucleic acid molecules that comprise the sequences of each of the plurality of nucleic acid targets; and / or ii. the number of first molecular labels having distinct sequences, second molecular labels having distinct sequences, or combinations thereof, associated with extended barcoded nucleic acid molecules of the plurality of extended barcoded nucleic acid molecules that comprise the sequences of each of the plurality of nucleic acid targets; The method may further comprise determining the copy number of each of the plurality of nucleic acid targets in the sample based on:

8. The method of claim 1, a. the sequence of each of the plurality of nucleic acid targets comprises a subsequence of each of the plurality of nucleic acid targets; and / or b. the sequence of the nucleic acid target within the plurality of barcoded nucleic acid molecules comprises a subsequence of the nucleic acid target; and / or c. the nucleic acid target comprises mRNA, and / or d. the sample comprises a single cell, optionally comprising an immune cell, and optionally comprising a B cell or a T cell, optionally wherein the single cell comprises a circulating tumor cell; and / or e. the sample comprises a plurality of cells, a plurality of single cells, a tissue, a tumor sample, or any combination thereof, optionally wherein the single cells comprise circulating tumor cells; and / or f. the first universal sequence of each oligonucleotide barcode of the first plurality of oligonucleotide barcodes is 5' to the first molecular label and the first target binding region; and / or the second universal sequence of each oligonucleotide barcode of the second plurality of oligonucleotide barcodes is 5' to the second molecular label and / or the second target binding region; The above method.

9. The method of claim 1, a. the method comprises amplifying a plurality of barcoded nucleic acid molecules using an amplification primer and a primer comprising a first universal sequence or a portion thereof, thereby generating a first plurality of single-labeled nucleic acid molecules comprising a sequence or a portion thereof of a nucleic acid target; determining the copy number of the nucleic acid target in the sample comprises determining the copy number of the nucleic acid target in the sample based on the number of first molecular labels having distinct sequences associated with the first plurality of single-labeled nucleic acid molecules or products thereof; Optionally, i. the amplification primer may contain a fourth universal sequence, and / or ii. the amplification primers may be target-specific primers; Optionally, the target-specific primers may specifically hybridize to an immune receptor, a constant region of an immune receptor, a variable region of an immune receptor, a diversity region of an immune receptor, and / or a junction between the variable and diversity regions of an immune receptor; Further optionally, the immune receptor may be a T cell receptor (TCR) and / or a B cell receptor (BCR) receptor; Optionally, the TCR may comprise a TCR alpha chain, a TCR beta chain, a TCR gamma chain, a TCR delta chain, or any combination thereof, and the BCR receptor may comprise a BCR heavy chain and / or a BCR light chain; and / or b. the method comprises amplifying the plurality of extended barcoded nucleic acid molecules using an amplification primer and a primer comprising a second universal sequence or a portion thereof, thereby generating a second plurality of single-labeled nucleic acid molecules comprising the sequence or a portion thereof of the nucleic acid target; determining the copy number of the nucleic acid target in the sample comprises determining the copy number of the nucleic acid target in the sample based on the number of second molecular labels having distinct sequences associated with the second plurality of single-labeled nucleic acid molecules or products thereof; Optionally, i. the amplification primer optionally comprises a fourth universal sequence; and / or ii. the amplification primers may be target-specific primers; Optionally, the target-specific primers may specifically hybridize to an immune receptor, a constant region of an immune receptor, a variable region of an immune receptor, a diversity region of an immune receptor, and / or a junction between the variable and diversity regions of an immune receptor; Further optionally, the immune receptor may be a T cell receptor (TCR) and / or a B cell receptor (BCR) receptor; Optionally, the TCR may comprise a TCR alpha chain, a TCR beta chain, a TCR gamma chain, a TCR delta chain, or any combination thereof, and the BCR receptor may comprise a BCR heavy chain and / or a BCR light chain; and / or c. the method hybridizing random primers to a plurality of barcoded nucleic acid molecules and extending the random primers to generate a first plurality of extension products, wherein the random primers comprise a third universal sequence or a complement thereof; amplifying the first plurality of extension products using a primer capable of hybridizing to a third universal sequence or its complement and a primer capable of hybridizing to the first universal sequence or its complement, thereby generating a third plurality of single-labeled nucleic acid molecules or products thereof; and / or d. determining the copy number of the nucleic acid target in the sample comprises determining the copy number of the nucleic acid target in the sample based on the number of first molecular labels having distinct sequences associated with a third plurality of single-labeled nucleic acid molecules or products thereof; and / or e. the method hybridizing random primers to the plurality of extended barcoded nucleic acid molecules and extending the random primers to generate a second plurality of extension products, wherein the random primers comprise a third universal sequence or a complement thereof; amplifying the second plurality of extension products using a primer capable of hybridizing to a third universal sequence or its complement and a primer capable of hybridizing to a second universal sequence or its complement, thereby generating a fourth plurality of single-labeled nucleic acid molecules or products thereof; and / or f. The method of any preceding claim, wherein determining the copy number of the nucleic acid target in the sample comprises determining the copy number of the nucleic acid target in the sample based on the number of second molecular labels having distinct sequences associated with a fourth plurality of single-labeled nucleic acid molecules or products thereof.

10. The method of claim 1, a. the first universal sequence, the second universal sequence, the third universal sequence, and / or the fourth universal sequence are the same; and / or b. the first universal sequence, the second universal sequence, the third universal sequence, and / or the fourth universal sequence are different; and / or c. the first universal sequence, the second universal sequence, the third universal sequence, and / or the fourth universal sequence comprise a binding site of a sequencing primer and / or a sequencing adapter, a complementary sequence thereof, and / or a portion thereof; Optionally, i. the sequencing adapter may comprise a P5 sequence, a P7 sequence, a complementary sequence thereof, and / or a portion thereof; and / or ii. The method of any one of the preceding claims, wherein the sequencing primer comprises a lead 1 sequencing primer, a lead 2 sequencing primer, a complementary sequence thereof, and / or a portion thereof.

11. The method of claim 1, a. obtaining sequence information of a plurality of barcoded nucleic acid molecules or products thereof, optionally wherein obtaining the sequence information can include attaching sequencing adapters to a plurality of extended barcoded nucleic acid molecules or products thereof; and / or b. obtaining sequence information of a plurality of extended barcoded nucleic acid molecules or products thereof, and optionally, obtaining sequence information may include attaching sequencing adapters to the extended barcoded nucleic acid molecules, the barcoded nucleic acid molecules, products thereof, or any combination thereof; and / or c. obtaining sequence information for one or more of the first, second, third, and fourth plurality of single-labeled nucleic acid molecules or products thereof; Optionally, i. obtaining sequence information may include attaching sequencing adaptors to one or more of the first, second, third, and fourth plurality of single-labeled nucleic acid molecules or products thereof; and / or ii. obtaining sequence information of one or more of the first, second, third, and fourth plurality of single-labeled nucleic acid molecules or products thereof may comprise obtaining sequencing data comprising a plurality of sequencing reads of one or more of the first, second, third, and fourth plurality of single-labeled nucleic acid molecules or products thereof; Each of the plurality of sequencing reads may comprise: (1) a cellular landmark sequence; (2) a molecular landmark sequence; and / or (3) a subsequence of a nucleic acid target; Optionally, the method may include aligning each of a plurality of sequencing reads of the nucleic acid target to a respective unique cell label sequence indicative of a single cell of the sample to generate an aligned sequence of the nucleic acid target; Further optionally, (1) the aligned sequences of the nucleic acid targets may comprise at least 50% of the cDNA sequence of the nucleic acid targets, at least 70% of the cDNA sequence of the nucleic acid targets, at least 90% of the cDNA sequence of the nucleic acid targets, or the full length of the cDNA sequence of the nucleic acid targets; and / or (2) the aligned sequences of the nucleic acid targets may include complementarity determining region 1 (CDR1), complementarity determining region 2 (CDR2), complementarity determining region 3 (CDR3), a variable region, the full length of a variable region, or a combination thereof; and / or (3) The aligned sequences of the nucleic acid targets may include a variable region, a diversity region, a junction of a variable region, a diversity region and / or a constant region, or any combination thereof; The above method.

12. The method of claim 1, the nucleic acid target is an immune receptor, optionally the immune receptor may comprise a BCR light chain, a BCR heavy chain, a TCR alpha chain, a TCR beta chain, a TCR gamma chain, a TCR delta chain, or any combination thereof; Optionally, a. the aligned sequences of the nucleic acid targets may include complementarity determining region 1 (CDR1), complementarity determining region 2 (CDR2), complementarity determining region 3 (CDR3), a variable region, the full length of a variable region, or a combination thereof; and / or b. The method above, wherein the aligned sequences of the nucleic acid targets may comprise variable regions, diversity regions, junctions of variable regions, diversity regions and / or constant regions, or any combination thereof.

13. The method according to claim 11 or 12, a. The step of obtaining sequence information comprises obtaining sequence information of a BCR light chain and a BCR heavy chain of a single cell; The sequence information of the BCR light chain and the BCR heavy chain includes the sequence of the complementarity determining region 1 (CDR1), CDR2, CDR3, or any combination thereof, of the BCR light chain and / or the BCR heavy chain; the method comprises pairing the BCR light chain and BCR heavy chain of the single cell based on the sequence information obtained; and / or the sample comprises a plurality of single cells, and the method comprises pairing the BCR light chain and BCR heavy chain of at least 50% of said single cells based on the sequence information obtained; and / or b. obtaining sequence information includes obtaining sequence information of a TCR alpha chain and a TCR beta chain of a single cell; The sequence information of the TCR alpha chain and the TCR beta chain includes the sequence of complementarity determining region 1 (CDR1), CDR2, CDR3, or any combination thereof, of the TCR alpha chain and / or the TCR beta chain; the method comprises pairing the TCR alpha and beta chains of the single cell based on the sequence information obtained; and / or the sample comprises a plurality of single cells, and the method comprises pairing the TCR alpha and beta chains of at least 50% of said single cells based on the sequence information obtained; and / or c. The step of obtaining sequence information includes obtaining sequence information of a TCR gamma chain and a TCR delta chain of a single cell; The sequence information of the TCR gamma chain and the TCR delta chain includes the sequence of complementarity determining region 1 (CDR1), CDR2, CDR3, or any combination thereof, of the TCR gamma chain and / or the TCR delta chain; the method comprises pairing the TCR gamma chain and the TCR delta chain of the single cell based on the sequence information obtained; and / or the sample comprises a plurality of single cells, and the method comprises pairing the TCR gamma chain and the TCR delta chain of at least 50% of the single cells based on the obtained sequence information. The above method.

14. The method of claim 1, a. the complement of the molecular beacon comprises the reverse complement of the molecular beacon or the complement of the molecular beacon; and / or b. the plurality of barcoded nucleic acid molecules comprises barcoded deoxyribonucleic acid (DNA) molecules, barcoded ribonucleic acid (RNA) molecules, or a combination thereof; and / or c. the nucleic acid target comprises a nucleic acid molecule, optionally the nucleic acid molecule may comprise ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNA containing a poly(A) tail, or any combination thereof, and further optionally the mRNA may encode an immune receptor; and / or d. the nucleic acid target comprises a cellular component binding reagent and / or the nucleic acid molecule is associated with a cellular component binding reagent, and optionally the method may further comprise the step of dissociating the nucleic acid molecule and the cellular component binding reagent; and / or e. at least 10 of the first and / or second plurality of oligonucleotide barcodes comprise different molecular label sequences; and / or f. each molecular label of the first and / or second plurality of oligonucleotide barcodes comprises at least six nucleotides; and / or g. the first and / or second plurality of oligonucleotide barcodes are associated with the solid support; Optionally, i. the first and / or second plurality of oligonucleotide barcodes associated with the same solid support may each comprise the same sample label, and optionally, each sample label of the first and / or second plurality of oligonucleotide barcodes may comprise at least 6 nucleotides; and / or ii. the first and / or second plurality of oligonucleotide barcodes may each comprise a cell label, and optionally, each cell label of the first and / or second plurality of oligonucleotide barcodes may comprise at least 6 nucleotides; and / or iii. Oligonucleotide barcodes in the first and / or second plurality of oligonucleotide barcodes associated with the same solid support may comprise the same cell label; and / or iv. the oligonucleotide barcodes of the first and / or second plurality of oligonucleotide barcodes associated with different solid supports may comprise different cell markers; and / or v. The solid support may comprise a synthetic particle, a planar surface, or a combination thereof; The above method.

15. The method of claim 1, a. the method comprises extending the oligonucleotide barcode in the presence of one or more of ethylene glycol, polyethylene glycol, 1,2-propanediol, dimethyl sulfoxide (DMSO), glycerol, formamide, 7-deaza-GTP, acetamide, tetramethylammonium chloride salt, betaine, or any combination thereof; and / or b. the sample comprises a single cell, and the method comprises associating a synthetic particle comprising a first and a second plurality of oligonucleotide barcodes with the single cell in the sample; Optionally, i. the method may comprise lysing the single cells after associating the synthetic particles with the single cells, and optionally lysing the single cells may comprise heating the sample, contacting the sample with a detergent, altering the pH of the sample, or any combination thereof; and / or ii. the synthetic particles and the single cells may be in the same compartment, optionally the compartment may be a well or a droplet; and / or iii. at least one oligonucleotide barcode of the first and / or second plurality of oligonucleotide barcodes may be immobilized or partially immobilized on a synthetic particle, or at least one oligonucleotide barcode of the first and / or second plurality of oligonucleotide barcodes may be encapsulated or partially encapsulated within a synthetic particle; and / or iv. The synthetic particles may be disintegrable, and optionally may be disintegrable hydrogel particles; the synthetic particles may comprise beads, optionally the beads may be sepharose beads, streptavidin beads, agarose beads, magnetic beads, conjugated beads, protein A conjugated beads, protein G conjugated beads, protein A / G conjugated beads, protein L conjugated beads, oligo(dT) conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads or any combination thereof; and / or The synthetic particles may comprise a material selected from the group consisting of polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic material, ceramic, plastic, glass, methylstyrene, acrylic polymer, titanium, latex, sepharose, cellulose, nylon, silicone, and any combination thereof; and / or v. each oligonucleotide barcode of the first and / or second plurality of oligonucleotide barcodes may comprise a linker functional group; The synthetic particles may include a solid support functional group, and The support functional group and the linker functional group may be associated with each other, and optionally the linker functional group and the support functional group may each be selected from the group consisting of C6, biotin, streptavidin, a primary amine, an aldehyde, a ketone, and any combination thereof; The above method.