Methods and compositions for combinatorial barcoding

Combinatorial barcoding methods using unique barcode subunit and linker arrays address the challenge of low throughput in nucleic acid tagging, enabling efficient and accurate sequencing of samples and single cells.

JP7893625B2Inactive Publication Date: 2026-07-22BECTON DICKINSON & CO
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
BECTON DICKINSON & CO
Filing Date
2022-03-01
Publication Date
2026-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Nucleic acid barcoding methods face challenges in generating a vast repertoire of barcodes for sample tagging due to low throughput of individual tagging with sample-specific barcodes before combination with other barcoded samples becomes possible.

Method used

Combinatorial barcoding methods involving sets of unique barcode subunit arrays and linker arrays or their complements are used to generate a large number of unique combinatorial barcodes, which are concatenated to associate each sample or target with a unique barcode, enabling efficient pooling and sequencing.

Benefits of technology

The method enables the generation of a vast repertoire of barcodes, facilitating high-throughput nucleic acid sequencing and characterization of samples, including single cells, with improved efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007893625000001
    Figure 0007893625000001
  • Figure 0007893625000002
    Figure 0007893625000002
  • Figure 0007893625000003
    Figure 0007893625000003
Patent Text Reader

Abstract

The present disclosure provides compositions, methods, and kits for generating sets of combinatorial barcodes and their use for barcoding samples such as single cells and genomic DNA fragments. Some embodiments disclosed herein provide a composition comprising a set of component barcodes for generating a set of combinatorial barcodes, which may include, for example, n x m unique component barcodes, where n and m are integers, each of the component barcodes comprising one of the n unique barcode subunit sequences and one or two linker sequences or their complements, and the component barcodes are configured to connect to each other via the one or two linker sequences or their complements to generate the set of combinatorial barcodes.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Incorporation by reference of priority application This application claims priority under U.S. Provisional Patent Application No. 62 / 140,360, filed March 30, 2015, and U.S. Provisional Patent Application No. 62 / 152,644, filed April 24, 2015. The contents of these applications are expressly incorporated by reference in their entirety in this application.

[0002] Reference to array list This application is filed together with an electronic sequence listing. The sequence listing is provided as a file named BDCRI-013WO_SEQLISTING.TXT, created on March 28, 2016, with a size of 695 bytes. The electronic information of the sequence listing is incorporated herein by reference in its entirety. [Background technology]

[0003] Miniaturization and parallel processing of individual reactions have led to dramatic cost reductions and throughput increases in modern scientific experiments and measurements. Nucleic acid barcoding can involve methods of identifying and tagging nucleic acids in each sample using sequence strings (also known as barcode sequences) to add informational content to the sequence. Barcoding may involve separation in chemical (barcode) space and may or may not rely on physical isolation. DNA barcoding can allow for a pooling of various sequences. DNA barcoding can be difficult to implement due to the low throughput of individual tagging with sample-specific barcodes before combination with other barcoded samples becomes possible. New methods are needed to generate a vast repertoire of barcodes for sample tagging. [Overview of the project] [Means for solving the problem]

[0004] Some embodiments disclosed herein are compositions comprising a set of component barcodes for generating a set of combinatorial barcodes, comprising n×m unique component barcodes (where n and m are positive integers), each of the component barcodes comprising one of n unique barcode subunit arrays and one or two linker arrays or their complements, and the component barcodes being configured to be connected to each other via one or two linker arrays or their complements to generate a set of combinatorial barcodes. In some embodiments, each of the sets of n unique barcode subunit arrays comprises a molecular label, a cell label, a dimensional label, a universal label, or any combination thereof. In some embodiments, the total number of different linker arrays is m - 1 or m - 2. In some embodiments, each of the component barcodes comprises a barcode subunit array-linker, a complement of the linker-barcode subunit array-linker, or a complement of the linker-barcode subunit array. In some embodiments, each of the sets of combinatorial barcodes has the formula: barcode subunit array a -linker 1-barcode subunit array b -linker 2-…barcode subunit array c -linker m-1 -barcode subunit array d and comprises an oligonucleotide. In some embodiments, the oligonucleotide comprises a target-specific region. In some embodiments, the target-specific region is oligo dT. In some embodiments, each of the sets of combinatorial barcodes has the formula: barcode subunit array a -linker 1-barcode subunit array b -linker 2-…-barcode subunit array c and a first oligonucleotide, and a barcode subunit array d -…-linker 3-barcode subunit array e -linker m-2 -barcode subunit arrayf A second oligonucleotide containing , and . In some embodiments, each of the first and second oligonucleotides contains a target-specific region. In some embodiments, one of the target-specific regions is oligo dT. In some embodiments, each of the one or two linker sequences is different. In some embodiments, m is an integer from 2 to 10. In some embodiments, m is an integer from 2 to 4. In some embodiments, n is an integer from 4 to 100. In some embodiments, n=24. In some embodiments, n unique barcode subunit sequences are of the same length. In some embodiments, at least two of the n unique barcode subunit sequences have different lengths. In some embodiments, the linker sequences are of the same length. In some embodiments, the linker sequences have different lengths. In some embodiments, the set of combinatorial barcodes is n m The set of unique combinatorial barcodes has a number of unique combinatorial barcodes or less. In some embodiments, the set of combinatorial barcodes has a number of unique combinatorial barcodes of at least 100,000. In some embodiments, the set of combinatorial barcodes has a number of unique combinatorial barcodes of at least 200,000. In some embodiments, the set of combinatorial barcodes has a number of unique combinatorial barcodes of at least 300,000. In some embodiments, the set of combinatorial barcodes has a number of unique combinatorial barcodes of at least 400,000. In some embodiments, the set of combinatorial barcodes has a number of unique combinatorial barcodes of at least 1,000,000.

[0005] Some embodiments disclosed herein provide a method for barcoded multiple partitions, comprising the steps of: introducing m component barcodes into each of the multiple partitions, wherein m component barcodes are selected from a set of n × m different component barcodes (where n and m are positive integers), and each component barcode comprises one of n unique barcode subunit sequences and one or two linker sequences or their complements; and concatenating the m component barcodes to generate a combinatorial barcode, thereby associating each of the multiple partitions with a unique combinatorial barcode comprising m component barcodes. In some embodiments, each of the multiple partitions contains a sample. In some embodiments, the sample is a DNA sample. In some embodiments, the sample is an RNA sample. In some embodiments, the sample is a single cell. In some embodiments, the total number of different linker sequences is m-1 or m-2. In some embodiments, the method includes the step of hybridizing one or two of the m component barcodes to a target in the sample. In some embodiments, the method includes the step of extending one or two of m-type component barcodes hybridized to a target, thereby labeling the target with a combinatorial barcode. In some embodiments, the component barcodes are introduced in each of several partitions using an inkjet printer.

[0006] Some embodiments disclosed herein provide a method for barcoding multiple DNA targets in multiple partitions, wherein m component barcodes are selected from a set of n × m different component barcodes (where n and m are positive integers), and each component barcode comprises one of n unique barcode subunit sequences and one or two linker sequences or their complements, the method comprising the steps of depositing the m component barcodes onto each of the multiple partitions containing the DNA targets, and concatenating the m component barcodes to generate a combinatorial barcode, thereby associating each DNA target in the multiple partitions with a unique combinatorial barcode comprising the m component barcodes. In some embodiments, the total number of different linker sequences is m-1 or m-2. In some embodiments, the method includes the step of hybridizing one or two of the m component barcodes to the DNA targets. In some embodiments, the method includes extending one or two m-type component barcodes hybridized to a DNA target to produce an extension product containing a combinatorial barcode. In some embodiments, the method includes amplifying the extension product. In some embodiments, the amplification step includes isothermal multistrand substitution amplification. In some embodiments, the method includes pooling the amplified products of multiple partitions. In some embodiments, the method includes sequencing the pooled amplified products. In some embodiments, the method includes assembling the sequence using the combinatorial barcode. In some embodiments, the method includes generating a haplotype using the assembled sequence. In some embodiments, the haplotype is generated using overlapping sequencing reads covering at least 100 kb.

[0007] Some embodiments disclosed herein provide a single-cell sequencing method comprising the steps of: depositing the m component barcodes onto each of a plurality of partitions containing single cells, under the condition that m component barcodes are selected from a set of n × m different component barcodes (where n and m are positive integers), and each component barcode includes one of n unique barcode subunit sequences and one or two linker sequences or their complements; and concatenating the m component barcodes to generate a combinatorial barcode, thereby associating the single-cell targets in each of the plurality of partitions with a unique combinatorial barcode containing the m component barcodes. In some embodiments, the total number of different linker sequences is m-1 or m-2. In some embodiments, the target is DNA. In some embodiments, the target is RNA. In some embodiments, the method includes the step of hybridizing one or two of the m component barcodes to a target. In some embodiments, the method includes the step of extending one or two m-type component barcodes hybridized to a target to produce an extension product containing a combinatorial barcode. In some embodiments, the method includes the step of amplifying the extension product. In some embodiments, the method includes the step of pooling the amplified products of multiple partitions. In some embodiments, the method includes the step of sequencing the pooled amplified products. In some embodiments, the method includes the step of assembling a sequence from a single cell using the combinatorial barcode. In some embodiments, the method includes the step of characterizing a single cell in large-scale parallel.

[0008] Several embodiments disclosed herein provide a method for affixing a spatial barcode to a target in a sample, comprising the steps of: providing a plurality of oligonucleotides immobilized on a substrate; contacting the sample with the plurality of oligonucleotides immobilized on the substrate; hybridizing the target with a probe that specifically binds to the target; capturing an image indicating the location of the target; capturing an image of the sample; and correlating the image indicating the location of the target with the image of the sample to affix a spatial barcode to the target in the sample. In some embodiments, the oligonucleotides immobilized on the substrate are oligo-dT primers. In some embodiments, the oligonucleotides immobilized on the substrate are random primers. In some embodiments, the oligonucleotides immobilized on the substrate are target-specific primers. In some embodiments, the method includes the step of reverse transcribing the target to produce cDNA. In some embodiments, the method includes the step of adding a dATP homopolymer to the cDNA. In some embodiments, the method includes the step of amplifying the cDNA using bridge amplification. In some embodiments, the substrate is part of a flow cell. In some embodiments, the method includes the step of dissolving the sample before or after the step of contacting the sample with the plurality of oligonucleotides immobilized on the substrate. In some embodiments, the step of dissolving the sample includes heating the sample, contacting the sample with a surfactant, changing the pH of the sample, or any combination thereof. In some embodiments, the sample includes tissue slices, cell monolayers, fixed cells, tissue sections, or any combination thereof. In some embodiments, the sample includes multiple cells. In some embodiments, the multiple cells include one or more cell types. In some embodiments, at least one of the one or more cell types is brain cells, cardiac cells, cancer cells, circulating tumor cells, organocytes, epithelial cells, metastatic cells, benign cells, primary cells, circulating cells, or any combination thereof. In some embodiments, the method includes the step of amplifying a target to produce multiple amplicons immobilized on a substrate.In some embodiments, the targets include ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNA containing a poly(A) tail, or any combination thereof. In some embodiments, the method includes the step of affixing spatial barcodes to multiple targets. In some embodiments, the probe is 18 to 100 nt in length. In some embodiments, the probe includes a fluorescent label. In some embodiments, the image indicating the location of the target is a fluorescent image. In some embodiments, the sample includes immunohistochemical staining, progressive staining, hematoxylin-eosin staining, or a combination thereof. In some embodiments, the solid carrier includes a polymer, matrix, hydrogel, needle array device, antibody, or any combination thereof.

[0009] Several embodiments disclosed herein provide a method for single-cell sequencing, comprising the steps of: depositing a first oligonucleotide comprising a first barcode and a first primer selected from a first set of 40 unique barcodes into a plurality of partitions containing single cells; depositing a second oligonucleotide comprising a second barcode and a second primer selected from a second set of 40 unique barcodes into a plurality of partitions containing single cells; contacting the first oligonucleotide with a single-cell target; extending the first oligonucleotide to produce a first chain comprising the first barcode; contacting the second oligonucleotide with the first chain; and extending the second oligonucleotide to produce a second chain comprising the first barcode and the second barcode, thereby labeling the single-cell target with the first barcode and the second barcode. In some embodiments, the method includes a step of amplifying the second chain to produce an amplification product. In some embodiments, the method includes a step of pooling the amplification products from the plurality of partitions. In some embodiments, the method includes a step of sequencing pooled amplification products. In some embodiments, the method includes a step of assembling a single-cell sequence using a first barcode and a second barcode. In some embodiments, the method includes a step of characterizing single cells in large-scale parallel configuration. In some embodiments, multiple partitions are composed of 1536 microwell plates. In some embodiments, the target is mRNA. In some embodiments, the first primer is an oligo-dT primer. In some embodiments, the target is DNA.

[0010] Some embodiments disclosed herein provide compositions comprising a set of combinatorial barcodes, wherein each combinatorial barcode in the set comprises m barcode subunit arrays selected from a set of n unique barcode subunit arrays (where n and m are positive integers) and m-1 or m-2 linker arrays, the m barcode subunit arrays being connected to each other via m-1 or m-2 linker arrays, and the total number of unique combinatorial barcodes in the set of combinatorial barcodes is greater than 384. In some embodiments, each of the n unique barcode subunit arrays comprises molecular labels, cell labels, dimensional labels, universal labels, or any combination thereof. In some embodiments, the total number of unique combinatorial barcodes in the set of combinatorial barcodes is n mThe following applies: In some embodiments, each set of combinatorial barcodes includes an oligonucleotide comprising the formula: barcode subunit sequence a-linker 1-barcode subunit sequence b-linker 2-...barcode subunit sequence c-linker m-1-barcode subunit sequence d. In some embodiments, each set of combinatorial barcodes includes a first oligonucleotide comprising the formula: barcode subunit sequence a-linker 1-barcode subunit sequence b-linker 2-...-barcode subunit sequence c, and a second oligonucleotide comprising barcode subunit sequence d-...-linker 3-barcode subunit sequence e-linker m-2-barcode subunit sequence f. In some embodiments, each set of combinatorial barcodes includes one or two target-specific regions. In some embodiments, one of the one or two target-specific regions is oligo dT. In some embodiments, each of the m-1 or m-2 linker sequences is different. In some embodiments, m is an integer from 2 to 10. In some embodiments, m is an integer from 2 to 4. In some embodiments, n is an integer between 4 and 100. In some embodiments, n = 24. In some embodiments, n unique barcode subunit sequences are of the same length. In some embodiments, n unique barcode subunit sequences have different lengths. In some embodiments, m-1 or m-2 linker sequences are of the same length. In some embodiments, m-1 or m-2 linker sequences have different lengths. In some embodiments, a set of combinatorial barcodes has at least 100,000 unique combinatorial barcodes. In some embodiments, a set of combinatorial barcodes has at least 200,000 unique combinatorial barcodes. In some embodiments, a set of combinatorial barcodes has at least 300,000 unique combinatorial barcodes. In some embodiments, a set of combinatorial barcodes has at least 400,000 unique combinatorial barcodes.In some embodiments, the set of combinatorial barcodes has at least 1,000,000 unique combinatorial barcodes. In some embodiments, each of the set of combinatorial barcodes is bound to a solid carrier. In some embodiments, each of the set of combinatorial barcodes is used to label a partition. In some embodiments, the partition is part of a microwell array having more than 10,000 microwells. In some embodiments, each microwell in the microwell array contains a different combinatorial barcode from the set of combinatorial barcodes.

[0011] In one embodiment, the present disclosure relates to a composition comprising a set of reagent barcodes, wherein each reagent barcode m has a length of n, the number of different reagent barcodes is n × m, and the possible number of different combinatorial barcodes is n m The present invention provides a composition that is as follows: In some embodiments, different combinatorial barcodes are combinations of reagent barcodes. In some embodiments, different combinatorial barcodes differ in at least one nucleotide. In some embodiments, different combinatorial barcodes differ in at least two nucleotides. In some embodiments, n m / (n×m) is at least 1:2000. In some embodiments, n m / (n×m) is at least 1:3000. In some embodiments, n m / (n×m) is at least 1:4000. In some embodiments, the number of different combinatorial barcodes is at least 100,000. In some embodiments, the number of different combinatorial barcodes is at least 200,000. In some embodiments, the number of different combinatorial barcodes is at least 300,000. In some embodiments, the number of different combinatorial barcodes is at least 400,000. In some embodiments, the number of different combinatorial barcodes is 331,776. In some embodiments, m is at least 2. In some embodiments, m is at least 3. In some embodiments, m is at least 4. In some embodiments, m is 4. In some embodiments, n is 5 to 50. In some embodiments, n is 24. In some embodiments, the reagent barcodes in the set of reagent barcodes include different lengths. In some embodiments, the first barcode is linked to the second barcode in the set of barcodes via a linker. In some embodiments, the first barcode includes a target-specific region. In some embodiments, the target-specific region hybridizes to the sense strand of the target polynucleotide. In some embodiments, the combinatorial barcode is bipertite. In some embodiments, the combinatorial barcode is tripartite. In some embodiments, the combinatorial code is partly located on the 3' end of the target polynucleotide and partly on the 5' end of the target polynucleotide. In some embodiments, the target-specific region hybridizes to the antisense strand of the target polynucleotide. In some embodiments, the composition further comprises the target polynucleotide.

[0012] In one embodiment, the present disclosure is a method for barcoded samples, wherein each reagent barcode m has a length of n, the number of different reagent barcodes is n × m, and the number of possible different combinatorial barcodes is n mThe present invention provides a method comprising the steps of: contacting a set of reagent barcodes with a plurality of partitions under certain conditions; associating the barcode reagent with a target polynucleotide; and labeling the target polynucleotide with the barcode reagent under conditions in which each target polynucleotide contains a different combinatorial barcode. In some embodiments, the target polynucleotide is DNA. In some embodiments, the target polynucleotide is genomic DNA. In some embodiments, the target polynucleotide is RNA. In some embodiments, the target polynucleotide is a chromosome fragment. In some embodiments, the target polynucleotide is at least 100 kilobases long. In some embodiments, the target polynucleotide is from a single cell.

[0013] In one embodiment, the disclosure provides a method for isolating a plurality of cells, comprising the steps of: non-Poissonally isolating single cells of the plurality of cells into a single well of a substrate under conditions that the well contains two or more combinatorial barcode reagents; and generating combinatorial barcoded nucleic acids by labeling the nucleic acids of the cells with the combinatorial barcode reagents. In some embodiments, the isolation step includes the step of distributing the plurality of cells such that at least 20% of the wells of the substrate contain single cells. In some embodiments, the isolation step includes the step of distributing the plurality of cells such that at least 40% of the wells of the substrate contain single cells. In some embodiments, the isolation step includes the step of distributing the plurality of cells such that at least 60% of the wells of the substrate contain single cells. In some embodiments, the isolation step includes the step of distributing the plurality of cells such that at least 80% of the wells of the substrate contain single cells. In some embodiments, the isolation step includes the step of distributing the plurality of cells such that 100% of the wells of the substrate contain single cells. In some embodiments, the isolation step is non-random. In some embodiments, the isolation step is performed by an isolation device. In some embodiments, the isolation device includes a device selected from the group consisting of a flow cytometer, a needle array, and a microinjector, or any combination thereof. In some embodiments, the substrate includes at least 500 wells. In some embodiments, the substrate includes at least 1,000 wells. In some embodiments, the substrate includes at least 10,000 wells. In some embodiments, the substrate includes 96, 384, or 1536 wells. In some embodiments, one combinatorial barcode reagent includes a subunit code section. In some embodiments, this combinatorial barcode reagent further includes a linker adjacent to the subunit code section. In some embodiments, the linker is located at 5', 3', or both 5' and 3' of the subunit code section.In some embodiments, linkers of different combinatorial barcode reagents are configured to hybridize and integrate to produce a concatenated combinatorial barcode reagent. In some embodiments, the concatenated barcode reagent corresponds to a combinatorial barcode. In some embodiments, the combinatorial barcode is bipertite. In some embodiments, the combinatorial barcode is partially on the 3' end of the combinatorial barcode-attached nucleic acid and partially on the 5' end of the combinatorial barcode-attached nucleic acid. In some embodiments, the subunit code section is 4 to 30 nucleotides long. In some embodiments, the subunit code section is 6 nucleotides long. In some embodiments, each subunit code section of the combinatorial barcode reagent contains a unique subunit code sequence. In some embodiments, one combinatorial barcode reagent contains a target-specific region. In some embodiments, only one of the combinatorial barcode reagents contains a target-specific region. In some embodiments, only two of the combinatorial barcode reagents contain target-specific regions. In some embodiments, the target-specific region is selected from the group consisting of random multimers, oligo-dTs, or gene-specific sequences. In some embodiments, the target-specific region is adapted to hybridize to the sense strand of the target nucleic acid. In some embodiments, the target-specific region is adapted to hybridize to the antisense strand of the target nucleic acid. In some embodiments, the labeling step includes a step of hybridizing the combinatorial barcode reagent to the nucleic acid. In some embodiments, the hybridizing step further includes a step of hybridizing the combinatorial barcode reagents to each other via a cognitive linker. In some embodiments, the method further includes a step of extending the combinatorial barcode reagent. In some embodiments, the extension step includes a step of primer extension of the combinatorial barcode reagent. In some embodiments, the extension step includes a step of generating a transcript containing the combinatorial barcode reagent sequence and the nucleic acid.In some embodiments, the substrate wells contain different combinations of combinatorial barcode reagents. In some embodiments, the different combinations of combinatorial barcode reagents correspond to different combinatorial barcodes. In some embodiments, the method further includes a step of pooling combinatorial barcoded nucleic acids after the labeling step. In some embodiments, the method further includes a step of amplifying the combinatorial barcoded nucleic acids, thereby generating combinatorial barcoded amplicons. In some embodiments, the amplification step includes multistrand substitution. In some embodiments, the amplification step includes PCR. In some embodiments, the amplification step is performed using primers that hybridize to the sequence of the combinatorial barcode reagents. In some embodiments, the amplification step is performed using gene-specific primers and primers that hybridize to the sequence of the combinatorial barcode reagents. In some embodiments, the method further includes a step of determining the sequence of the combinatorial barcoded amplicons. In some embodiments, the determination step includes determining a portion of the combinatorial barcode sequence and a portion of the nucleic acid sequence of the combinatorial barcoded amplicon. In some embodiments, the method further includes lysing cells before the labeling step. In some embodiments, the nucleic acid is RNA. In some embodiments, the nucleic acid is mRNA. In some embodiments, the nucleic acid is DNA. In some embodiments, the nucleic acid is genomic DNA. In some embodiments, the cells include cells selected from the group consisting of human cells, mammalian cells, rat cells, pig cells, mouse cells, fly cells, worm cells, invertebrate cells, vertebrate cells, fungal cells, bacterial cells, and plant cells, or any combination thereof. In some embodiments, the cells include tumor cells. In some embodiments, the cells include disease cells. In some embodiments, the number of combinatorial barcode reagents in the wells is at least two. In some embodiments, the number of combinatorial barcode reagents in the wells is at least three.In some embodiments, the number of combinatorial barcode reagents in the wells is at least 4.

[0014] In one embodiment, the disclosure provides a kit comprising a set of combinatorial barcode reagents, wherein only one of the combinatorial barcode reagents comprises a target-specific sequence, and the barcodes of the set of combinatorial barcode reagents comprise linkers such that the combinatorial barcode reagents overlap via linkers, thereby generating combinatorial barcodes. In some embodiments, the combinatorial barcode comprises a combination of combinatorial barcode reagents. In some embodiments, the sequences of the combinatorial barcode differ by at least one nucleotide. In some embodiments, the sequences of the combinatorial barcode differ by at least two nucleotides. In some embodiments, the number of combinatorial barcodes generated from the set of combinatorial barcode reagents is at least 100,000. In some embodiments, the number of combinatorial barcodes generated from the set of combinatorial barcode reagents is at least 200,000. In some embodiments, the number of combinatorial barcodes generated from the set of combinatorial barcode reagents is at least 300,000. In some embodiments, the number of combinatorial barcodes generated from a set of combinatorial barcode reagents is at least 400,000. In some embodiments, the number of combinatorial barcodes generated from a set of combinatorial barcode reagents is 331,776. In some embodiments, one combinatorial barcode reagent includes a subunit code section. In some embodiments, this combinatorial barcode reagent further includes a linker adjacent to the subunit code section. In some embodiments, the linker is located at 5', 3', or both 5' and 3' of the subunit code section. In some embodiments, the linkers of different combinatorial barcode reagents are configured to hybridize and integrate to produce a concatenated combinatorial barcode reagent. In some embodiments, the concatenated barcode reagent includes a combinatorial barcode.In some embodiments, the subunit code section is 5 to 35 nucleotides long. In some embodiments, the subunit code section is 6 nucleotides long. In some embodiments, each subunit code section of the combinatorial barcode contains a unique subunit code sequence. In some embodiments, each reagent barcode m has a subunit code length of n, and n. m / (n×m) is at least 2000. In some embodiments, n m / (n×m) is at least 3000. In some embodiments, n m / (n×m) is at least 4000. In some embodiments, n×m is the number of different reagent barcodes. In some embodiments, n m is the number of possible different combinatorial barcodes. In some embodiments, one combinatorial barcode reagent includes a target-specific region. In some embodiments, the target-specific region is selected from the group consisting of random multimers, oligo dTs, or gene-specific sequences. In some embodiments, the target-specific region is adapted to hybridize to the sense strand of the target nucleic acid. In some embodiments, the target-specific region is adapted to hybridize to the antisense strand of the target nucleic acid. In some embodiments, the combinatorial barcode is bipertite. In some embodiments, a portion of the combinatorial barcode is located partly on the 3' end of the target polynucleotide and partly on the 5' end of the target polynucleotide. In some embodiments, the kit further includes a substrate. In some embodiments, the substrate includes at least 1,000 microwells. In some embodiments, the microwells include at least two reagent barcodes from a set of reagent barcodes. In some embodiments, the kit further includes instructions for use. In some embodiments, the kit further includes a buffer. In some embodiments, the kit further includes gene-specific primers. In some embodiments, the kit further includes amplification primers.

[0015] In one embodiment, the present disclosure provides a method for whole-genome sequencing comprising: fragmenting a nucleic acid chromosome into one or more chromosomal fragments; isolating one or more chromosomal fragments in wells of a substrate containing two or more reagent barcodes; generating combinatorial barcoded fragments by amplifying the chromosomal fragments together with the reagent barcodes; and performing whole-genome sequencing by determining the sequence of the combinatorial barcoded fragments. In some embodiments, the chromosomal fragments are at least 100 kilobases long. In some embodiments, the chromosomal fragments are 50 to 300 kilobases long. In some embodiments, the fragmentation step is 1 × 10⁻¹⁶ 6 ~1 × 10 7 Fragments are generated. In some embodiments, the isolation step includes isolating 1 to 5 fragments in a well. In some embodiments, the 1 to 5 fragments do not overlap. In some embodiments, the amplification step includes multistrand substitution. In some embodiments, the amplification step is isothermal. In some embodiments, the method further includes pooling the combinatorial barcoded fragments before the determination step. In some embodiments, the determination step includes sequencing at least a portion of the combinatorial barcodes and at least a portion of the chromosome fragments. In some embodiments, the determination step includes grouping the sequenced reads into several groups based on the combinatorial barcode sequences. In some embodiments, the groups correspond to wells of the substrate. In some embodiments, the method further includes assembling contigs from the chromosome fragments. In some embodiments, the method further includes mapping the contigs to maternal or paternal chromosomes. In some embodiments, the mapping step determines the haplotype phasing of the chromosome fragments. In certain embodiments, for example, the following are provided: (Item 1) A composition comprising a set of component barcodes for generating a set of combinatorial barcodes, n × m types of unique component barcodes, where n and m are positive integers. Includes, Each of the aforementioned component barcodes is, One of n unique barcode subunit arrays, One or two linker sequences or their complements, Includes, A composition in which the component barcodes are configured to be linked together via the one or two linker arrays or their complements to generate a set of combinatorial barcodes. (Item 2) The composition according to item 1, wherein each of the n sets of unique barcode subunit sequences comprises a molecular label, a cellular label, a dimensional label, a universal label, or any combination thereof. (Item 3) The composition according to item 1 or 2, wherein the total number of different linker sequences is m-1 or m-2. (Item 4) Each of the aforementioned component barcodes is, a) Barcode subunit array - linker, b) Linker complement - barcode subunit array - linker, or c) Linker complement - barcode subunit array, A composition containing any one of items 1 to 3. (Item 5) Each of the combinatorial barcodes in the set is represented by the formula: barcode subunit array. a - Linker 1 - Barcode subunit arrangement b - Linker 2 -...Barcode subunit arrangement c - Linker m-1 - Barcode subunit arrangement d A composition according to any one of items 1 to 4, comprising an oligonucleotide containing the above. (Item 6) The composition according to item 5, wherein the oligonucleotide comprises a target-specific region. (Item 7) The composition according to item 6, wherein the target-specific region is oligo dT. (Item 8) Each of the combinatorial barcode sets is Formula: Barcode subunit array a - Linker 1 - Barcode subunit arrangement b - Linker 2 -…-Barcode subunit arrangement c A first alkyl group containing, Barcode subunit array d -...-Linker 3 - Barcode subunit arrangement e - Linker m-2 - Barcode subunit arrangement f A second alkyl group containing, A composition containing any one of items 1 to 4. (Item 9) The composition according to item 8, wherein each of the first oligonucleotide and the second oligonucleotide comprises a target-specific region. (Item 10) The composition according to item 9, wherein one of the target-specific regions is oligo dT. (Item 11) The composition according to any one of items 1 to 10, wherein one of each of the one or two linker sequences is different. (Item 12) A composition according to any one of items 1 to 11, wherein m is an integer from 2 to 10. (Item 13) A composition according to any one of items 1 to 11, wherein m is an integer between 2 and 4. (Item 14) A composition according to any one of items 1 to 13, wherein n is an integer between 4 and 100. (Item 15) A composition according to any one of items 1 to 13, wherein n=24. (Item 16) The composition according to any one of items 1 to 15, wherein the n unique barcode subunit sequences are of the same length. (Item 17) The composition according to any one of items 1 to 15, wherein at least two of the n unique barcode subunit sequences have different lengths. (Item 18) The composition according to any one of items 1 to 17, wherein the linker sequences are of the same length. (Item 19) The composition according to any one of items 1 to 17, wherein the linker array has different lengths. (Item 20) A set of combinatorial barcodes m A composition according to any one of items 1 to 19, having a unique combinatorial barcode of type 1 or less. (Item 21) The composition according to any one of items 1 to 19, wherein the set of combinatorial barcodes has at least 100,000 unique combinatorial barcodes. (Item 22) The composition according to any one of items 1 to 19, wherein the set of combinatorial barcodes has at least 200,000 unique combinatorial barcodes. (Item 23) The composition according to any one of items 1 to 19, wherein the set of combinatorial barcodes has at least 300,000 unique combinatorial barcodes. (Item 24) The composition according to any one of items 1 to 19, wherein the set of combinatorial barcodes has at least 400,000 unique combinatorial barcodes. (Item 25) The composition according to any one of items 1 to 19, wherein the set of combinatorial barcodes has at least 1,000,000 unique combinatorial barcodes. (Item 26) A method for attaching barcodes to multiple partitions, m types of component barcodes are selected from a set of n × m types of different component barcodes, where n and m are positive integers, and each of the component barcodes is One of n unique barcode subunit arrays, One or two linker sequences or their complements, Conditions including, The process involves introducing m types of component barcodes to each of multiple partitions, The process of generating a combinatorial barcode by connecting the aforementioned m types of component barcodes, Includes, A method wherein each of the plurality of partitions is associated with a unique combinatorial barcode containing the m types of component barcodes. (Item 27) The method according to item 26, wherein each of the aforementioned partitions contains a sample. (Item 28) The method according to item 27, wherein the sample is a DNA sample. (Item 29) The method described in item 27, wherein the sample is an RNA sample. (Item 30) The method described in item 27, wherein the sample is a single cell. (Item 31) The method according to any one of items 26-30, wherein the total number of different linker sequences is m-1 or m-2. (Item 32) The method according to any one of items 27 to 31, comprising the step of hybridizing one or two of the m types of component barcodes to a target of the sample. (Item 33) The method according to item 32, comprising the step of extending one or two of the m-type component barcodes hybridized to the target, thereby labeling the target with the combinatorial barcode. (Item 34) The method according to any one of items 26 to 33, wherein the component barcode is introduced into each of the plurality of partitions using an inkjet printer. (Item 35) A method for affixing barcodes to multiple DNA targets in multiple partitions, m types of component barcodes are selected from a set of n × m types of different component barcodes, where n and m are positive integers, and each of the component barcodes is One of n unique barcode subunit arrays, One or two linker sequences or their complements, Conditions including, A step of depositing the m types of component barcodes on each of the plurality of partitions containing a DNA target, The process of generating a combinatorial barcode by connecting the aforementioned m types of component barcodes, Includes, A method wherein each DNA target in the plurality of partitions is associated with a unique combinatorial barcode containing the m types of component barcodes. (Item 36) The method according to item 35, wherein the total number of different linker sequences is either m-1 or m-2. (Item 37) The method according to item 35 or 36, comprising the step of hybridizing one or two of the m types of component barcodes to the DNA target. (Item 38) The method according to item 37, comprising the step of extending one or two of the m-type component barcodes hybridized to the DNA target to produce an extension product comprising the combinatorial barcode. (Item 39) The method according to item 38, comprising the step of amplifying the extension product. (Item 40) The method according to item 39, wherein the step of amplifying the extension product includes isothermal multi-chain substitution amplification. (Item 41) The method according to item 39 or 40, comprising the step of pooling the amplification products of the plurality of partitions. (Item 42) The method according to item 41, comprising the step of sequencing the pooled amplification products. (Item 43) The method according to item 42, comprising the step of assembling the sequence using the combinatorial barcode. (Item 44) The method according to item 43, comprising the step of generating a haplotype using the assembled sequence. (Item 45) The method according to item 44, wherein the haplotype is generated by a step of overlapping sequencing reads that cover at least 100 kb. (Item 46) A single-cell sequencing method, m types of component barcodes are selected from a set of n × m types of different component barcodes, where n and m are positive integers, and each of the component barcodes is One of n unique barcode subunit arrays, One or two linker sequences or their complements, Conditions including, A step of depositing the m-type component barcodes on each of a plurality of partitions containing a single cell, The process of generating a combinatorial barcode by connecting the aforementioned m types of component barcodes, Includes, A method wherein each single cell target in the plurality of partitions is associated with a unique combinatorial barcode containing the m types of component barcodes. (Item 47) The method described in item 46, wherein the total number of different linker sequences is either m-1 or m-2. (Item 48) The method according to item 46 or 47, wherein the target is DNA. (Item 49) The method according to item 46 or 47, wherein the target is RNA. (Item 50) The method according to any one of items 46 to 49, comprising the step of hybridizing one or two of the m types of component barcodes to the target. (Item 51) The method according to item 50, comprising the step of stretching one or two of the m-type component barcodes hybridized to the target to produce a stretched product including the combinatorial barcode. (Item 52) The method according to item 51, comprising the step of amplifying the extension product. (Item 53) The method according to item 52, comprising the step of pooling amplification products from the plurality of partitions. (Item 54) The method according to item 53, comprising the step of sequencing the pooled amplification products. (Item 55) The method according to item 54, comprising the step of assembling a sequence from a single cell using the combinatorial barcode. (Item 56) The method described in item 55, which includes a step of large-scale parallel characterization of single cells. (Item 57) A method for attaching spatial barcodes to targets in a sample, A step of providing multiple oligonucleotides immobilized on a substrate, A step of bringing a sample into contact with a plurality of oligonucleotides immobilized on the substrate, and a step of hybridizing the target with a probe that specifically binds to the target, A step of capturing an image showing the location of the target, The process of capturing an image of the aforementioned sample, A step of correlating an image showing the location of the target with an image of the sample and attaching a spatial barcode to the target in the sample, A method that includes this. (Item 58) The method according to item 57, wherein the oligonucleotide immobilized on the substrate is an oligo dT primer. (Item 59) The method according to item 57, wherein the oligonucleotide immobilized on the substrate is a random primer. (Item 60) The method according to item 57, wherein the oligonucleotide immobilized on the substrate is a target-specific primer. (Item 61) The method according to any one of items 57 to 60, comprising the step of reverse transcribing the target to generate cDNA. (Item 62) The method according to item 61, comprising the step of adding a dATP homopolymer to the cDNA. (Item 63) The method according to item 62, comprising the step of amplifying cDNA using bridge amplification. (Item 64) The method according to any one of items 57 to 63, wherein the substrate is part of a flow cell. (Item 65) The method according to any one of items 57 to 64, comprising the step of dissolving the sample before or after the step of contacting the sample with a plurality of oligonucleotides immobilized on the substrate. (Item 66) The method according to item 65, wherein the step of dissolving the sample includes a step of heating the sample, a step of contacting the sample with a surfactant, a step of changing the pH of the sample, or any combination thereof. (Item 67) The method according to any one of items 57 to 66, wherein the sample comprises a tissue slice, a cell monolayer, fixed cells, a tissue section, or any combination thereof. (Item 68) The method according to any one of items 57 to 66, wherein the sample comprises multiple cells. (Item 69) The method according to item 68, wherein the plurality of cells include one or more cell types. (Item 70) The method according to item 69, wherein at least one of the one or more cell types is a brain cell, cardiac cell, cancer cell, circulating tumor cell, organocyte, epithelial cell, metastatic cell, benign cell, primary cell, circulating cell, or any combination thereof. (Item 71) The method according to any one of items 57 to 70, comprising the step of amplifying the target to generate a plurality of amplicons fixed on the substrate. (Item 72) The method according to any one of items 57 to 70, wherein the target includes ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNA containing a poly(A) tail, or any combination thereof. (Item 73) The method according to any one of items 57 to 72, comprising the step of affixing spatial barcodes to multiple targets. (Item 74) The method according to any one of items 57 to 73, wherein the probe has a length of 18 to 100 nt. (Item 75) The method according to any one of items 57 to 74, wherein the probe includes a fluorescent label. (Item 76) The method according to item 75, wherein the image showing the location of the target is a fluorescence image. (Item 77) The method according to any one of items 57 to 76, wherein the sample is subjected to immunohistochemical staining, progressive staining, hematoxylin-eosin staining, or a combination thereof. (Item 78) The method according to any one of items 57 to 77, wherein the solid carrier comprises a polymer, a matrix, a hydrogel, a needle array device, an antibody, or any combination thereof. (Item 79) A single-cell sequencing method, A step of depositing a first alkyl group, comprising a first barcode selected from a first set of 40 unique barcodes and a first primer, into multiple partitions containing single cells, A step of depositing a second oligonucleotide, comprising a second barcode selected from a second set of 40 unique barcodes and a second primer, into a plurality of partitions containing the single cell, A step of bringing the first oligonucleotide into contact with the target of the single cell, A step of extending the first oligonucleotide to generate a first chain containing the first barcode, A step of bringing the second oligonucleotide into contact with the first chain, A step of extending the second oligonucleotide to generate a second chain containing the first barcode and the second barcode, Includes, A method wherein the single cell target is labeled with the first barcode and the second barcode. (Item 80) The method according to item 79, further comprising the step of amplifying the second chain to produce an amplification product. (Item 81) The method according to item 80, comprising the step of pooling amplification products from the plurality of partitions. (Item 82) The method according to item 80 or 81, comprising the step of sequencing the pooled amplification products. (Item 83) The method according to item 82, comprising the step of assembling a sequence from a single cell using the first barcode and the second barcode. (Item 84) The method described in item 83, which includes a step of large-scale parallel characterization of single cells. (Item 85) The method according to any one of items 79 to 84, wherein the plurality of partitions are comprised of a 1536 microwell plate. (Item 86) The method according to any one of items 79 to 85, wherein the target is mRNA. (Item 87) The method according to item 86, wherein the first primer is an oligo-dT primer. (Item 88) The method according to any one of items 79 to 85, wherein the target is DNA. (Item 89) A composition comprising a set of combinatorial barcodes, Each combinatorial barcode in the aforementioned set of combinatorial barcodes is A set of n unique barcode subunit arrays, each consisting of m barcode subunit arrays selected from a set of n unique barcode subunit arrays, where n and m are positive integers, m-1 or m-2 linker sequences, Includes, A composition in which the m barcode subunit arrays are connected to each other via m-1 or m-2 linker arrays, and the total number of unique combinatorial barcodes in the set of combinatorial barcodes is greater than 384. (Item 90) The composition according to item 89, wherein each of the set of n unique barcode subunit sequences comprises a molecular label, a cellular label, a dimensional label, a universal label, or any combination thereof. (Item 91) The total number of unique combinatorial barcodes in the set of combinatorial barcodes is n m The compositions described in item 89 or 90, which are as follows: (Item 92) Each of the aforementioned sets of combinatorial barcodes is represented by the formula: barcode subunit array. a - Linker 1 - Barcode subunit arrangement b - Linker 2 -...Barcode subunit arrangement c - Linker m-1 - Barcode subunit arrangement d A composition according to any one of items 89 to 91, comprising an oligonucleotide containing the above. (Item 93) Each of the aforementioned sets of combinatorial barcodes is: Formula: Barcode subunit array a - Linker 1 - Barcode subunit arrangement b - Linker 2 -…-Barcode subunit arrangement c A first alkyl group containing, Barcode subunit array d -...-Linker 3 - Barcode subunit arrangement e - Linkerm-2 - Barcode subunit arrangement f A second alkyl group containing, A composition containing any one of items 89 to 91. (Item 94) The composition according to any one of items 89 to 93, wherein each of the combinatorial barcodes comprises one or two target-specific regions. (Item 95) The composition according to item 94, wherein one of the one or two target-specific regions is oligo dT. (Item 96) The composition according to any one of items 89 to 95, wherein each of the m-1 or m-2 linker sequences is different. (Item 97) A composition according to any one of items 89 to 96, wherein m is an integer from 2 to 10. (Item 98) A composition according to any one of items 89 to 96, wherein m is an integer between 2 and 4. (Item 99) A composition according to any one of items 89 to 98, wherein n is an integer between 4 and 100. (Item 100) A composition according to any one of items 89 to 98, wherein n=24. (Item 101) The composition according to any one of items 89 to 100, wherein the n unique barcode subunit sequences are of the same length. (Item 102) The composition according to any one of items 89 to 100, wherein the n unique barcode subunit sequences have different lengths. (Item 103) The composition according to any one of items 89 to 102, wherein the m-1 or m-2 linker sequences are of the same length. (Item 104) The composition according to any one of items 89 to 102, wherein the m-1 or m-2 linker arrays have different lengths. (Item 105) The composition according to any one of items 89 to 104, wherein the set of combinatorial barcodes has at least 100,000 unique combinatorial barcodes. (Item 106) The composition according to any one of items 89 to 104, wherein the set of combinatorial barcodes has at least 200,000 unique combinatorial barcodes. (Item 107) The composition according to any one of items 89 to 104, wherein the set of combinatorial barcodes has at least 300,000 unique combinatorial barcodes. (Item 108) The composition according to any one of items 89 to 104, wherein the set of combinatorial barcodes has at least 400,000 unique combinatorial barcodes. (Item 109) The composition according to any one of items 89 to 104, wherein the set of combinatorial barcodes has at least 1,000,000 unique combinatorial barcodes. (Item 110) The composition according to any one of items 89 to 109, wherein one of each set of combinatorial barcodes is bound to a solid carrier. (Item 111) The composition according to any one of items 89 to 110, wherein one of each set of combinatorial barcodes is used to label a partition. (Item 112) The composition according to item 111, wherein the partition is part of a microwell array having more than 10,000 microwells. (Item 113) The composition according to item 112, wherein each microwell in the microwell array contains a combinatorial barcode different from the set of combinatorial barcodes.

[0016] Novel features of the present invention are specifically described in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by referring to the following detailed description and accompanying drawings illustrating exemplary embodiments in which the principles of the present invention are utilized. [Brief explanation of the drawing]

[0017] [Figure 1] Exemplary embodiments of the combinatorial barcode reagents of this disclosure are shown. [Figure 2] This disclosure illustrates exemplary embodiments of concatenation of combinatorial barcode reagents. [Figure 3] Exemplary embodiments of the method of combinatorial barcoding of nucleic acids of this disclosure are shown. [Figure 4] This disclosure illustrates an exemplary embodiment of a method for generating contigs using the combinatorial barcode reagents of this disclosure. [Figure 5] This disclosure provides exemplary embodiments of a method for genotyping single cells using the combinatorial barcode reagents of this disclosure. [Figure 6] An exemplary embodiment of a method combining combinatorial barcoding and probabilistic barcoding is shown. [Figure 7] This disclosure illustrates an exemplary embodiment of the Viper-Tite Combinatorial Barcode Method. [Figure 8] An exemplary embodiment of a combinatorial barcode fixed on a solid carrier is shown. [Figure 9] An exemplary embodiment of a method for affixing a spatial barcode to a target in a sample is shown. [Modes for carrying out the invention]

[0018] This disclosure provides compositions and methods for generating a repertoire of various barcoding reagents using only a small set of component barcoding reagents. The methods of this disclosure provide a simple molecular biology step for combinatorial pairing of component barcoding reagents. In some cases, parallel processing of hundreds of thousands to millions of samples is possible. For example, the number of samples that can be processed in parallel may be at least 100, 1,000, 5,000, 10,000, 50,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, or 900,000. In some embodiments, the number of samples that can be processed in parallel may be at most 100, 1,000, 5,000, 10,000, 50,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, or 900,000. In some embodiments, it is possible to process at least 1 million, 2 million, 3 million, 4 million, 5 million, 6 million, 7 million, 8 million, 9 million, or 10 million or more samples in parallel. In some embodiments, parallel processing of at most 1 million, 2 million, 3 million, 4 million, 5 million, 6 million, 7 million, 8 million, 9 million, or 10 million or more samples is possible.

[0019] This disclosure provides methods, compositions, kits, and systems for combinatorial barcoding of nucleic acids using combinatorial barcode reagents. As shown in Figure 1, the combinatorial barcode reagents 105 / 110 / 115 / 120 of this disclosure may include a target-specific region, a subunit code section, and a linker region, or any combination thereof. The subunit code section XXXXXX may have a subunit code sequence. Different subunit code sections may have different subunit code sequences. As shown in Figure 1, the combinatorial barcode reagents 105, 110, 115, and 120 are concatenated integrally via the linker region. The concatenated combinatorial barcode reagents can be extended (e.g., by primer extension) to produce a transcript 125 containing the combinatorial barcode reagent sequence and a target polynucleotide. The combinatorial barcode reagents can reduce the number of reagents required to produce a large repertoire of barcodes.

[0020] definition Unless otherwise specified, all technical terms used herein have the same meaning as those generally understood by those skilled in the art of the field to which this disclosure pertains. Where used herein and in the appended claims, the singular forms "a," "an," and "the" encompass multiple reference words unless otherwise explicitly stated in the context. Wherever "or" is used herein, unless otherwise explicitly stated, it is intended to encompass "and / or."

[0021] As used herein, the terms “associated” or “associated with” may mean that two or more species are identifiable as being co-located at some point in time. Association may mean that two or more species are in similar containers. Association may be an informatic association, in which case, for example, digital information about two or more species is stored and that information can be used to determine that one or more of these species are co-located. Association may also be a physical association, in which case two or more associated species are “tethered,” “joined,” or “fixed” to each other or to a common solid or semi-solid surface. Association may mean covalent or non-covalent means for attaching a label to a solid or semi-solid support, such as a bead. Association may include hybridization of a target and a label.

[0022] The terms “component barcode,” “reagent barcode,” and “combinatorial barcode reagent” are used synonymously to mean a polynucleotide sequence containing a subunit code sequence, which can be used in combination with one or more other combinatorial barcode reagents to generate a combinatorial barcode. For example, a combinatorial barcode reagent can generate a combinatorial barcode by concatenating it via a linker.

[0023] As used herein, the term “combinatorial barcode” means a polynucleotide sequence containing the sequences of one or more combinatorial barcode reagents.

[0024] As used herein, the term “complementary” refers to the ability of two nucleotides to precisely pair. For example, two nucleic acids are considered complementary at a given position if a nucleotide at a given position in one nucleic acid can form a hydrogen bond with a nucleotide in another nucleic acid. Complementarity between two single-stranded nucleic acid molecules can be “partial” if only some of the nucleotides are bonded, or it can be complete if complementarity exists throughout the entire single-stranded molecule. If a first nucleotide sequence is complementary to a second nucleotide sequence, then the first nucleotide sequence can be said to be the “complement” of the second sequence. If a first nucleotide sequence is complementary to the reverse of the second sequence (i.e., the nucleotides are in the reverse order), then the first nucleotide sequence can be said to be the “reverse complement” of the second sequence. As used herein, the terms “complement,” “complementary,” and “reverse complement” can be used synonymously. It is understood from this disclosure that if one molecule can hybridize with another molecule, that molecule may be a complement to the molecule it is hybridizing with.

[0025] As used herein, the term “digital counting” means a method for estimating the number of target molecules in a sample. Digital counting may include the step of determining the number of unique labels associated with a target in the sample. This probabilistic method transforms the problem of counting molecules from a problem of locating and identifying identical molecules into a series of yes / no digital problems concerning the detection of a given set of labels.

[0026] As used herein, the term “first universal label” means a label that is universal to the barcodes of this disclosure. The first universal label may be a sequencing primer binding site (for example, a lead primer binding site, i.e., for an Illumina sequencer).

[0027] As used herein, the term “label” means a nucleic acid code associated with a target in a sample. A label may, for example, be a nucleic acid label. A label may be a label that is amplified in whole or in part. A label may be a label that is sequenceable in whole or in part. A label may be a portion of a native nucleic acid that can be individually identified. A label may be a known sequence. A label may include a conjugation of nucleic acid sequences (for example, a conjugation of a native sequence and a non-native sequence). As used herein, the term “label” may be used synonymously with the terms “index,” “tag,” or “labeled tag.” A label is information-transmitting. For example, in various embodiments, a label can be used to determine sample identity, sample source, cell identity, and / or target.

[0028] As used herein, the term “non-exhausted reservoir” means a pool of probability barcodes composed of a wide variety of labels. A non-exhausted reservoir may contain a large number of different probability barcodes such that, if the non-exhausted reservoir is associated with a pool of targets, each target is likely to be associated with a unique probability barcode. The uniqueness of each labeled target molecule can be determined by the statistics of random selection and depends on the number of copies of identical target molecules in the collection compared to the diversity of labels. The size of the resulting set of labeled target molecules can be determined by the probabilistic nature of the barcoding process, and then analysis of the number of detected probability barcodes allows for the calculation of the number of target molecules present in the original collection or sample. If the ratio of the number of copies of present target molecules to the number of unique probability barcodes is low, the labeled target molecules are highly unique (i.e., the probability that two or more target molecules are labeled with one given label is very low).

[0029] As used herein, “nucleic acid” means a polynucleotide sequence or a fragment thereof. Nucleic acids may contain nucleotides. Nucleic acids may be exogenous or endogenous to cells. Nucleic acids may exist in a cell-free environment. Nucleic acids may be genes or fragments thereof. Nucleic acids may be DNA. Nucleic acids may be RNA. Nucleic acids may contain one or more analogs (e.g., modified backbone, sugar, or nucleic acid base). Some examples of analogs, but not limited to, include 5-bromouracil, peptide nucleic acids, xeno nucleic acids, morpholino forms, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein conjugated to sugar), thiol-containing nucleotides, biotin-conjugated nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, cuosin, and waiosin. The terms "nucleic acid," "polynucleotide," "targeted polynucleotide," and "targeted nucleic acid" can be used synonymously.

[0030] Nucleic acids may contain one or more modifications (e.g., base modifications, skeletal modifications) to provide nucleic acids with new or improved characteristics (e.g., improved stability). Nucleic acids may contain nucleic acid affinity tags. Nucleosides can be base-sugar combinations. The base portion of a nucleoside can be a heterocyclic base. The two most common classes of such heterocyclic bases are purines and pyrimidines. Nucleotides can be nucleosides further containing a phosphate group covalently bonded to the sugar portion of the nucleoside. In nucleosides containing pentofuranosyl sugars, the phosphate group can be bonded to the 2', 3', or 5' hydroxyl portion of the sugar. When forming nucleic acids, the phosphate group can covalently bond adjacent nucleosides to each other to form linear polymer compounds. Subsequently, each end of these linear polymer compounds can be further linked to form cyclic compounds. However, linear compounds are generally preferred. In addition, linear compounds may have internal nucleotide base complementarity and can fold to produce fully double-stranded or partially double-stranded compounds. Within nucleic acids, phosphate groups are usually referable as forming the internucleoside skeleton of the nucleic acid. The bond or skeleton of nucleic acids can be a 3'→5' phosphodiester bond.

[0031] Nucleic acids may contain modified skeletons and / or modified internucleoside bonds. Modified skeletons may include those that retain a phosphorus atom in the skeleton and those that do not. Preferred modified nucleic acid skeletons containing a phosphorus atom may include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkyl phosphotriesters, methyl and other alkyl phosphonates such as 3'-alkylene phosphonates and 5'-alkylene phosphonates, chiral phosphonates, phosphinates, phosphoramidates such as 3'-aminophosphoramides and aminoalkylphosphoramides, phosphorodiamidates, thionophosphoramides, thionoalkyl phosphonates, thionoalkyl phosphotriesters, selenophosphates, and boranophosphates, which usually have 3'-5', 2'-5' bond analogues, as well as those with reverse polarity, where one or more internucleotide bonds are 3'→3', 5'→5', or 2'→2' bonds.

[0032] Nucleic acids may contain polynucleotide skeletons formed by short-chain alkyl or cycloalkyl nucleoside bonds, mixed heteroatoms and alkyl or cycloalkyl nucleoside bonds, or one or more short-chain heteroatoms or heterocycle nucleoside bonds. These may include morpholino bonds (partially formed from the sugar moiety of a nucleoside), siloxane skeletons, sulfides, sulfoxides, and sulfone skeletons, formacetyl and thioformacetyl skeletons, methyleneformacetyl and thioformacetyl skeletons, riboacetyl skeletons, alkene-containing skeletons, sulfamate skeletons, methyleneimino and methylenehydrazino skeletons, sulfonate and sulfonamide skeletons, amide skeletons, and others having mixed N, O, S, and CH2 components.

[0033] Nucleic acids can contain nucleic acid mimetic compounds. The term "mimetic" includes, for example, polynucleotides in which only the furanose ring or both the furanose ring and the internucleotide bond are replaced with non-furanose groups, and substitution of only the furanose ring can be referred to as a sugar surrogate. Heterocyclic base moieties or modified heterocyclic base moieties are retainable for hybridization with suitable target nucleic acids. One such nucleic acid may be a peptide nucleic acid (PNA). In PNAs, the sugar backbone of the polynucleotide is replaceable with an amide-containing backbone, particularly an aminoethylglycine backbone. The nucleotide is retainable and is directly or indirectly bonded to the aza nitrogen atom of the amide moiety of the backbone. The backbone in a PNA compound may contain two or more bonded aminoethylglycine units that give the PNA an amide-containing backbone. The heterocyclic base moiety is directly or indirectly bondable to the aza nitrogen atom of the amide moiety of the backbone.

[0034] Nucleic acids may contain a morpholino skeletal structure. For example, a nucleic acid may contain a six-membered morpholino ring instead of a ribose ring. In some of these embodiments, the phosphodiester bond can be replaced by an internucleoside bond of a phosphorodiamidate or other non-phosphodiester.

[0035] Nucleic acids may contain linked morpholino units (i.e., morpholino nucleic acids) having heterocyclic bases linked to a morpholino ring. The binding group can link morpholino monomer units in morpholino nucleic acids. Nonionic morpholino oligomeric compounds may have fewer undesirable interactions with cellular proteins. Morpholino polynucleotides can be nonionic mimics of nucleic acids. Various compounds within the morpholino class can be linked using different binding groups. A further class of polynucleotide mimetic can be referred to as cyclohexenyl nucleic acids (CeNA). The furanose ring normally present in nucleic acid molecules can be replaced with a cyclohexenyl ring. CeNA DMT-protected phosphoramidite monomers can be prepared and used for the synthesis of oligomeric compounds using phosphoramidite chemistry. Incorporation of CeNA monomers into nucleic acid chains can increase the stability of DNA / RNA hybrids. CeNA oligoadenylates can form complexes with nucleic acid complements that exhibit stability similar to that of natural complexes. Further modifications may include locked nucleic acids (LNAs) in which a 2'-hydroxyl group is bonded to the 4' carbon atom of the sugar ring to form a 2'-C,4'-C-oxymethylene bond, thereby forming a bicyclic sugar moiety. The bond may be a methylene (-CH2-) group (wherein n is 1 or 2) bridging the 2' oxygen atom and the 4' carbon atom. LNAs and LNA analogs may exhibit very high double-strand thermal stability (Tm = +3 to +10°C), stability against 3'-exonuclease degradation, and good solubility with complementary nucleic acids.

[0036] Nucleic acids may also, in some embodiments, involve modifications or substitutions of nucleic acid bases (often simply referred to as “bases”). As used herein, “unmodified” or “natural” nucleic acid bases may include purine bases (e.g., adenine (A) and guanine (G)) and pyrimidine bases (e.g., thymine (T), cytosine (C), and uracil (U)). Modified nucleic acid bases include other synthetic and natural nucleic acid bases, such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine, and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl(-C=C-CH3)uracil and cytosine, as well as other alkynyl derivatives of pyrimidine bases, 6-azouracil, This may include cytosine, and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl, and other 8-substituted adenines and guanines, 5-halo especially 5-bromo, 5-trifluoromethyl, and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-aminoadenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine, and 3-deazaguanine and 3-deaadenine. Modified nucleic acid bases may include tricyclic pyrimidines, such as phenoxazinecytidine (1H-pyrimido(5,4-b)(1,4)benzoxazine-2(3H)-one), phenothiazinecytidine (1H-pyrimido(5,4-b)(1,4)benzothiadin-2(3H)-one), G-clamps such as substituted phenoxazinecytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazine-2(3H)-one), carbazolecytidine (2H-pyrimido(4,5-b)indole-2-one), and pyridoindolecytidine (H-pyrimido(3',':4,5)pyrrolo[2,3-d]pyrimidine-2-one).

[0037] As used herein, the term "sample" means a composition containing a target. Suitable samples for analysis by the methods, devices, and systems of this disclosure include cells, single cells, tissues, organs, or organisms.

[0038] As used herein, the terms “sampling device” or “device” mean a device capable of taking a section of a sample and / or placing the section onto a substrate. A sample device may mean, for example, a fluorescence-activated cell sorting (FACS) machine, a cell sorter, a biopsy needle, a biopsy device, a tissue sectioning device, a microfluidic device, a blade grid, and / or a microtome.

[0039] As used herein, the term “solid carrier” means a discrete solid or semi-solid surface on which multiple probability barcodes can be bound together. A solid carrier may include any type of solid, porous, or hollow sphere, ball, bearing, cylinder, or other similar construct made of plastic, ceramic, metal, or polymer material (e.g., hydrogel) on which nucleic acids can be immobilized (e.g., covalently or non-covalently). A solid carrier may include discrete particles that are spherical (e.g., microspheres) or non-spherical or irregular in shape, such as cubic, rectangular, pyramidal, cylindrical, conical, oblate, or disc-shaped. Multiple solid carriers arranged spaced apart in an array may not contain a substrate. The term “solid carrier” may be used synonymously with the term “beads.”

[0040] The term "solid carrier" may mean "substrate." A substrate may be a type of solid carrier. A substrate may mean a continuous solid or semi-solid surface on which the methods of this disclosure can be performed. A substrate may mean, for example, an array, cartridge, chip, device, and slide. As used herein, "solid carrier" and "substrate" may be used synonymously.

[0041] As used herein, the term "probability barcode" refers to a polynucleotide sequence that includes an identifier of the present disclosure. A probability barcode can be a polynucleotide sequence that can be used for probability barcoding. A probability barcode can quantify a target in a sample. A probability barcode can be used to control errors that can occur after associating an identifier with a target. For example, a probability barcode can evaluate amplification or sequencing errors. A probability barcode associated with a target can be referred to as a probability barcode target or a probability barcode tag target.

[0042] As used herein, the term "probability barcoding" refers to random labeling (e.g., barcoding) of nucleic acids. Probability barcoding can utilize a recursive Poisson strategy to associate an identifier with a target and quantify the associated identifier. As used herein, the term "probability barcoding" can be used synonymously with "probability labeling".

[0043] As used herein, the term "target" refers to a composition that can be associated with a probability barcode. Exemplary targets suitable for analysis by the methods, devices, and systems of the present disclosure include oligonucleotides, DNA, RNA, mRNA, microRNA, tRNA, and the like. A target can be single-stranded or double-stranded. In some embodiments, the target can be a protein. In some embodiments, the target can be a lipid.

[0044] The term "reverse transcriptase" means a group of enzymes having reverse transcriptase activity (i.e., catalyzing the synthesis of DNA from an RNA template). Generally, such enzymes include, but are not limited to, retroviral reverse transcriptase, retrotransposon reverse transcriptase, retroplasmid reverse transcriptase, retrone reverse transcriptase, bacterial reverse transcriptase, group II intron-derived reverse transcriptase, and their mutants, variants, or derivatives. Non-retroviral reverse transcriptases include non-LTR retrotransposon reverse transcriptase, retroplasmid reverse transcriptase, retrone reverse transcriptase, and group II intron reverse transcriptase. Examples of group II intron reverse transcriptases include Lactococcus lactis Ll.LtrB intron reverse transcriptase, Thermosynechococcus elongatus TeI4c intron reverse transcriptase, or Geobacillus stearothermophilus GsI-IIC intron reverse transcriptase. Other classes of reverse transcriptases may include many classes of non-retroviral reverse transcriptases (i.e., retrons, group II introns, and particularly diversity-generating retroelements).

[0045] The term "template switching" refers to the ability of a reverse transcriptase to switch an initial nucleic acid sequence template to the 3' end of a new nucleic acid sequence template that has little to no complementarity with the 3' end of the nucleic acid synthesized from the initial template. Nucleic acid copies of target polynucleotides can be prepared using template switching. Template switching enables the preparation of DNA copies using a reverse transcriptase that allows for the synthesis of a continuous product DNA in which an adapter sequence is directly attached to a target oligonucleotide sequence without ligation, by switching an initial nucleic acid sequence template to the 3' end of a new nucleic acid sequence template that has little to no complementarity with the 3' end of the DNA synthesized from the initial template. Template switching may involve ligation of the adapter, homopolymer tailing (e.g., polyadenylation), random primers, or oligonucleotides that can be associated by polymerase.

[0046] Probability barcode As disclosed herein, a probability barcode may be a polynucleotide sequence that can be used to probabilistically label (e.g., a barcode, tag) a target. A probability barcode may include one or more labels. Exemplary labels, but not limited to, include universal labels, cellular labels, molecular labels, sample labels, plate labels, spatial labels, pre-spatial labels, and any combination thereof. A probability barcode may include a 5' amine that can bind the probability barcode to a solid support. A probability barcode may include one or more of the following: universal labels, dimensional labels, spatial labels, cellular labels, and molecular labels. The universal label may be the 5'-side label. The molecular label may be the 3'-side label. The spatial, dimensional, and cellular labels may be in any order. In some cases, the universal, spatial, dimensional, dimensional, cellular, and molecular labels are in any order. A probability barcode may include a target-binding region. The target-binding region is interactable with a target in the sample (e.g., a target nucleic acid, RNA, mRNA, DNA). For example, the target binding region may contain an oligo-dT sequence that can interact with the poly(A) tail of mRNA. In some cases, the probabilistic barcode label (e.g., universal label, dimensional label, spatial label, cellular label, and molecular label) may separate 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides.

[0047] A probability barcode may include one or more universal labels. One or more universal labels may be identical on all probability barcodes in a set of probability barcodes (e.g., bound to a given solid carrier). In some embodiments, one or more universal labels may be identical on all probability barcodes bound to multiple beads. In some embodiments, the universal label includes a nucleic acid sequence that can hybridize to a sequencing primer. Sequencing primers can be used to sequence probability barcodes containing the universal label. Sequencing primers (e.g., universal sequencing primers) may include sequencing primers associated with a high-throughput sequencing platform. In some embodiments, the universal label may include a nucleic acid sequence that can hybridize to a PCR primer. In some embodiments, the universal label includes nucleic acid sequences that can hybridize to both sequencing primers and PCR primers. The nucleic acid sequence of the universal label that can hybridize to a sequencing primer or PCR primer may be referred to as a primer binding site. The universal label may include a sequence that can be used to initiate transcription of the probability barcode. A universal label may include a sequence that can be used to extend a probabilistic barcode or a region within a probabilistic barcode. A universal label may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides long or more, or at least approximately such a length. A universal label may contain at least about 10 nucleotides. A universal label may be at most about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides long or more. In some embodiments, a cleavable linker or modified nucleotide is part of the universal label sequence that enables the cleavage and removal of the probabilistic barcode from the carrier. As used herein, a universal label may be used synonymously with a “universal PCR primer.”

[0048] Probabilistic barcodes may include dimensional labels. Dimensional labels may include nucleic acid sequences that provide information about the dimension in which probabilistic labeling was performed. For example, a dimensional label can provide information about the time when the target was probabilistically barcoded. Dimensional labels can be associated with the time of probabilistic barcoding of a sample. Dimensional labels can be activated at the time of probabilistic labeling. Different dimensional labels can be activated at different time points. Dimensional labels provide information about the target, the group of targets, and / or the order in which the probabilistic barcodes were applied to the sample. For example, a cell population can be probabilistically barcoded during the G0 phase of the cell cycle. Cells can be pulsed again with probabilistic barcodes during the G1 phase of the cell cycle. Cells can be pulsed again with probabilistic barcodes during the S phase of the cell cycle, and so on. Each pulse (e.g., each phase of the cell cycle) may contain a different dimensional label. Thus, dimensional labels provide information about which phase of the cell cycle the target was labeled. Dimensional labels can scrutinize a wide variety of biological time points. Examples of biological time, though not limited to them, include the cell cycle, transcription (e.g., transcription initiation), and transcript degradation. Another example is the possibility of probabilistic labeling of samples (e.g., cells, cell populations) before and / or after drug treatment and / or therapy. Changes in the copy number of identifiable targets can be indicators of a sample's response to a drug and / or therapy.

[0049] In some embodiments, the dimensional label is activatable. An activatable dimensional label is, for example, activatable at a specific point in time. An activatable dimensional label is constitutively activatable (e.g., not switchable off). An activatable dimensional label is reversibly activatable (e.g., an activatable dimensional label can be switched on and off). A dimensional label is reversibly activatable at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more times. For example, a dimensional label can be activated by fluorescence, light, chemical events (e.g., cleavage, ligation of other molecules, addition of modifications (e.g., pegylation, SUMOylation, acetylation, methylation, demethylation, deacetylation)), photochemical events (e.g., photocaging), and introduction of non-native nucleotides.

[0050] The dimensional identifier may be identical for all probability barcodes bound to a given solid carrier (e.g., beads), but may differ for different solid carriers (e.g., beads). In some embodiments, at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100% of the probability barcodes on the same solid carrier include the same dimensional identifier. In some embodiments, at least 60% of the probability barcodes on the same solid carrier include the same dimensional identifier. In some embodiments, at least 95% of the probability barcodes on the same solid carrier include the same dimensional identifier.

[0051] Multiple solid carriers (e.g., beads) have 10 6A number of unique dimensional label sequences of a certain magnitude or greater may exist. A dimensional label may be, for example, 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides long or more, or roughly at least such a length. A dimensional label may be at most about 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 20, 15, 12, 10, 9, 8, 7, 6, 5, 4 nucleotides long or less or more. A dimensional label may be, for example, about 5 to about 200 nucleotides long or about 10 to about 150 nucleotides long. In some embodiments, a dimensional label is about 20 to about 125 nucleotides long.

[0052] Probabilistic barcodes may include spatial labels. Spatial labels may include nucleic acid sequences that provide information about the spatial orientation of the target molecule associated with the probabilistic barcode. Spatial labels can be associated with coordinates in a sample. The coordinates may be fixed coordinates. For example, the coordinates can be fixed relative to the substrate. Spatial labels may be relative to a two-dimensional or three-dimensional grid. The coordinates can be fixed relative to a landmark. The landmark is identifiable in space. The landmark may be an imaginable structure. The landmark may be a biological structure, such as an anatomical landmark. The landmark may be a cellular landmark (e.g., an organelle). The landmark may be a non-natural landmark, such as a structure with an identifiable identifier, such as a color code, barcode, magnetism, fluorescence, radioactivity, or unique size or shape. Spatial labels can be associated with physical partitions (e.g., wells, containers, or droplets). In some cases, multiple spatial labels are used together to code one or more locations in space.

[0053] A spatial identifier may be identical for all probability barcodes bound to a given solid carrier (e.g., beads), but may differ for different solid carriers (e.g., beads). In some embodiments, at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100% of probability barcodes on the same solid carrier include the same spatial identifier. In some embodiments, at least 60% of probability barcodes on the same solid carrier include the same spatial identifier. In some embodiments, at least 95% of probability barcodes on the same solid carrier include the same spatial identifier.

[0054] Multiple solid carriers (e.g., beads) have 10 6 A number of unique spatially labeled sequences of a certain magnitude or greater may exist. The spatial label may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides long or more, or roughly at least such a nucleotide length. In some embodiments, the spatial label may be at most about 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 20, 15, 12, 10, 9, 8, 7, 6, 5, 4 nucleotides long. The spatial label may be, for example, about 5 to about 200 nucleotides long. The spatial label may be, for example, about 10 to about 150 nucleotides long. The spatial label may be about 20 to about 125 nucleotides long.

[0055] Probability barcodes may include cell labels. Cell labels may include nucleic acid sequences that provide information for determining which target nucleic acids originated from which cells. In some embodiments, the cell labels are identical for all probability barcodes bound to a given solid carrier (e.g., beads), but different for different solid carriers (e.g., beads). In some embodiments, at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100% of probability barcodes on the same solid carrier include the same cell label. In some embodiments, at least 60% of probability barcodes on the same solid carrier include the same cell label. In some embodiments, at least 95% of probability barcodes on the same solid carrier include the same cell label.

[0056] Multiple solid carriers (e.g., beads) have 10 6 A number of unique cell-labeling sequences of a certain magnitude or greater may be present. The cell label may be at least about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides long or more. The cell label may be at most about 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 20, 15, 12, 10, 9, 8, 7, 6, 5, 4 nucleotides long or less or more. The cell label may be, for example, about 5 to about 200 nucleotides long. The cell label may be, for example, about 10 to about 150 nucleotides long. The cell label may be, for example, about 20 to about 125 nucleotides long.

[0057] The probabilistic barcode may include molecular labels. The molecular labels may include nucleic acid sequences that provide identification of information regarding a specific type of target nucleic acid species hybridized to the probabilistic barcode. The molecular labels may include nucleic acid sequences that provide a counter for a specific presence of a target nucleic acid species hybridized to the probabilistic barcode (e.g., a target binding region). In some embodiments, various sets of molecular labels are bound to a given solid carrier (e.g., beads). In some embodiments, approximately 10⁶ or more unique molecular label sequences can be bound to a given solid carrier (e.g., beads). In some embodiments, approximately 10⁵ or more unique molecular label sequences can be bound to a given solid carrier (e.g., beads). In some embodiments, approximately 10⁴ or more unique molecular label sequences can be bound to a given solid carrier (e.g., beads). In some embodiments, approximately 10³ or more unique molecular label sequences can be bound to a given solid carrier (e.g., beads). In some embodiments, approximately 10² or more unique molecular label sequences can be bound to a given solid carrier (e.g., beads). Molecular labels can be at least approximately 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides long or longer. Molecular labels can be at most approximately 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 20, 15, 12, 10, 9, 8, 7, 6, 5, 4 nucleotides long or less.

[0058] The probability barcode may include a target binding region. In some embodiments, the target binding region includes a nucleic acid sequence that specifically hybridizes to a target (e.g., a target nucleic acid, a target molecule, e.g., a cellular nucleic acid to be analyzed), e.g., a specific gene sequence. In some embodiments, the target binding region includes a nucleic acid sequence that can bind (e.g., hybridize) to a specific position on a particular target nucleic acid. In some embodiments, the target binding region includes a nucleic acid sequence that is capable of specific hybridization to a restriction site overhang (e.g., an EcoRI attachment end overhang). The probability barcode can then ligate to any nucleic acid molecule containing a sequence complementary to the restriction site overhang.

[0059] The probability barcode may include a target-binding region. The target-binding region is hybridizable to the target of interest. For example, the target-binding region may contain an oligo dT that is hybridizable to mRNA containing a polyadenylated end. The target-binding region may be gene-specific. For example, the target-binding region can be configured to hybridize to a specific region of the target. The target-binding region may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides long or longer, or at least such a nucleotide length. The target-binding region can be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides or longer. The target-binding region can be 5 to 30 nucleotides long. If a probability barcode contains a gene-specific target-binding region, the probability barcode can be referenced as a gene-specific probability barcode.

[0060] The target binding region may include a nonspecific target nucleic acid sequence. A nonspecific target nucleic acid sequence may mean a sequence that can bind to multiple target nucleic acids independently of a specific sequence of the target nucleic acid. For example, the target binding region may include a random multimer sequence or an oligo-dT sequence that hybridizes to the poly-A tail of an mRNA molecule. A random multimer sequence may be, for example, a sequence of random dimers, random trimers, random quadromers, random pentamers, random hexamers, random septamers, random octamers, random nonomers, random decamers, or higher-order random multimers of any length. In some embodiments, the target binding region is identical for all probability barcodes bound to a given bead. In some embodiments, the target binding regions of multiple probability barcodes bound to a given bead include two or more different target binding sequences. The target binding region may be 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides long or more, or roughly at least such a nucleotide length. In some embodiments, the target-binding region is at most about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides long or longer.

[0061] A probabilistic barcode may contain an orientation that can be used for orienting (e.g., alignment) the probabilistic barcode. A probabilistic barcode may contain portions for isoelectric focusing. Different probabilistic barcodes may contain different isoelectric focusing points. Once such a probabilistic barcode is introduced into a sample, the sample can undergo isoelectric focusing to orient the probabilistic barcode into a known form. Thus, the orientation can be used to create a known map of the probabilistic barcode in the sample. Exemplary orientations, but not limited to, include electrophoretic mobility (e.g., based on the size of the probabilistic barcode), isoelectric focus, spin, conductivity, and / or self-assembly. For example, a probabilistic barcode may contain a self-assembly orientation and be able to self-assemble to a specific orientation upon activation (e.g., nucleic acid nanostructure).

[0062] Probabilistic barcodes may include affinity. Spatial labeling may include affinity. Affinity may include chemical and / or biological components that facilitate the binding of probabilistic barcodes to other entities (e.g., cell receptors). For example, affinity may include antibodies. Antibodies may be specific to specific parts on a sample (e.g., receptors). Antibodies can guide probabilistic barcodes to specific cell types or molecules. Targets in specific cell types or molecules and / or their vicinity can be probabilistically labeled. Because antibodies can guide probabilistic barcodes to specific locations, affinity can also provide spatial information in addition to the nucleotide sequence of spatial labeling. Antibodies may be therapeutic antibodies. Antibodies may be monoclonal antibodies. Antibodies may be polyclonal antibodies. Antibodies may be humanized. Antibodies may be chimeric. Antibodies may be naked antibodies. Antibodies may be fusion antibodies.

[0063] Antibodies can be full-length immunoglobulin molecules (i.e., naturally occurring or formed by normal immunoglobulin gene fragment recombination processes) (e.g., IgG antibodies) or immunoactive (i.e., specifically binding) portions of immunoglobulin molecules, such as antibody fragments.

[0064] Antibodies can be antibody fragments. Antibody fragments can be parts of antibodies such as F(ab')2, Fab', Fab, Fv, and sFv. Antibody fragments can bind to the same antigen recognized by the full-length antibody. Antibody fragments may include isolated fragments consisting of the variable region of an antibody, such as "Fv" fragments consisting of heavy and light chain variable regions, and recombinant single-chain polypeptide molecules ("scFv proteins") in which the light and heavy chain variable regions are linked by a peptide linker. Exemplary antibodies include, but are not limited to, antibodies against cancer cells, antibodies against viruses, antibodies that bind to cell surface receptors (CD8, CD34, CD45), and therapeutic antibodies.

[0065] The cell label and / or any label of the present disclosure may further include a unique set of nucleic acid subsequences of a defined length, for example 7 nucleotides each (equivalent to the number of bits used in some Hamming error correction codes), designed to provide error correction capabilities. Hamming codes, like other error correction codes, are based on the principle of redundancy and can be constructed by adding redundant parity bits to data transferred through a noisy medium. Such error correction codes can encode a sample identifier together with redundant parity bits and "transfer" this sample identifier as a "codeword". Hamming codes may mean an arithmetic process that identifies a unique binary code based on a unique redundancy that allows correction of 1-bit errors. For example, Hamming codes can be matched to nucleic acid barcodes to screen for single nucleotide errors that occur during nucleic acid amplification. Thereby, identification of single nucleotide errors using Hamming codes may enable correction of nucleic acid barcodes.

[0066] Hamming codes can be represented by a subset of possible codewords selected based on the center of a multi-dimensional sphere (i.e., a hypersphere, for example) in a binary subspace. A 1-bit error can be corrected because it can fall within the hypersphere associated with a particular codeword. On the other hand, a 2-bit error not associated with a particular codeword can be detected but not corrected. Consider a first hypersphere centered at the coordinates (0,0,0) (i.e., using, for example, an x-y-z coordinate system). In this case, any 1-bit error can be corrected because it is included within a radius of 1 from the center coordinates. That is, for example, a 1-bit error having coordinates of (0,0,0), (0,1,0), (0,0,1), (1,0,0), or (1,1,0). Similarly, a second hypersphere can be constructed. In this case, a 1-bit error can be corrected because it is included within a radius of 1 from its center coordinates (1,1,1) (i.e., for example, (1,1,1), (1,0,1), (0,1,0), or (0,1,1).

[0067] In some embodiments, the length of the nucleic acid subsequence used to generate the error correction code can vary. For example, it may be at least 3 nucleotides long, at least 7 nucleotides long, at least 15 nucleotides long, or at least 31 nucleotides long. In some embodiments, nucleic acid subsequences of other lengths can be used to generate the error correction code.

[0068] If a probabilistic barcode contains two or more markers of a certain type (e.g., two or more cellular markers or two or more molecular markers), the markers may intersperse linker marker sequences. These linker marker sequences can be at least approximately 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides long or longer. In some cases, the linker marker sequences are 12 nucleotides long. The linker marker sequences can facilitate the synthesis of the probabilistic barcode. The linker markers may include error correction (e.g., Hamming) codes.

[0069] Combinatorial barcode Numerous different combinatorial barcodes can be used to label multiple samples, such as single cells or nucleic acid fragments, for parallel analysis such as sequencing. As disclosed herein, a large number of sets of combinatorial barcodes can be generated from a relatively small number of barcode subunit sequences, under conditions where barcode subunit sequences are connected in various combinations. In this way, it is possible to significantly increase the number of different combinatorial barcodes generated while keeping the increase in the number of different barcode subunit sequences, the number of barcode subunit sequences connected to each other in the combinatorial barcode, or both, relatively small.

[0070] Several embodiments disclosed herein provide compositions comprising a set of combinatorial barcodes. A combinatorial barcode comprises two or more barcode subunit sequences connected to each other via a linker sequence, a target nucleotide sequence, or both. Two or more barcode subunit sequences can form a combinatorial barcode by connecting to each other, for example, via hybridization and subsequent extension of one or more linker sequences. For example, other barcode subunit sequences can be incorporated and repeated by hybridizing a linker sequence bound to one barcode subunit sequence to a linker sequence bound to another barcode subunit sequence, and then using the other barcode subunit sequence as a template in an extension reaction. In some embodiments, it is possible to hybridize both of two barcode subunit sequences to a target nucleotide sequence, and then perform two extension reactions to incorporate the two barcode subunit sequences into the target nucleotide sequence.

[0071] In some embodiments, the number of barcode subunit arrays in a combinatorial barcode, i.e., the length of the combinatorial barcode, can be influenced by the design of the linker array. For example, to generate a combinatorial barcode with m barcode subunit arrays, m-1 or m-2 linker arrays may be used, where m is an integer ≥ 2. If n unique barcode subunit arrays are used to generate a combinatorial barcode with m barcode subunit arrays, the maximum number is n. mIt is possible to generate n unique combinatorial barcodes. In some embodiments, the combinatorial barcode may include an oligonucleotide comprising the formula: barcode subunit sequence a-linker 1-barcode subunit sequence b-linker 2-... barcode subunit sequence c-linker m-1-barcode subunit sequence d. In some embodiments, the combinatorial barcode may include a first oligonucleotide comprising the formula: barcode subunit sequence a-linker 1-barcode subunit sequence b-linker 2-...-barcode subunit sequence c, and a second oligonucleotide comprising barcode subunit sequence d-...-linker 3-barcode subunit sequence e-linker m-2-barcode subunit sequence f. The first and second oligonucleotides may be linked to each other by a target nucleotide sequence. In some embodiments, each of the barcode subunit sequences a, b, c, ... in the combinatorial barcode is selected from a set of n unique barcode subunit sequences. In some embodiments, some or all of the barcode subunit sequences a, b, c, ... in a combinatorial barcode may be identical. In some embodiments, some or all of the barcode subunit sequences a, b, c, ... in a combinatorial barcode may be different.

[0072] In some embodiments, the set of combinatorial barcodes includes at least 1,000, at least 10,000, at least 100,000, at least 200,000, at least 300,000, at least 400,000, at least 500,000, at least 1,000,000, at least 10,000,000, at least 100,000,000, and at least 1,000,000,000 or more unique combinatorial barcodes.

[0073] The combinatorial barcodes disclosed herein may comprise one or more probabilistic barcodes. For example, a combinatorial barcode may comprise one or more universal labels, cell labels, molecular labels, sample labels, plate labels, spatial labels, pre-spatial labels, or any combination thereof. In some embodiments, a probabilistic barcode such as a cell label, sample label, or spatial label may comprise two or more barcode subunit sequences, for example, at least two, at least three, at least four, at least five, or more barcode subunit sequences.

[0074] In some embodiments, combinatorial barcodes can be immobilized on solid carriers such as beads or microwells in a microwell array. As shown in Figure 8, solid carriers such as beads or microparticles can be coated with a first plurality of combinatorial barcodes. In some embodiments, a combinatorial barcode may include one or more universal sequences (US), cell labels, linkers, molecular labels, and target-specific regions, such as oligo(dT) sequences. In some embodiments, the solid carrier may include a second plurality of combinatorial barcodes having the same structure (not necessarily the same sequence) as the first plurality of combinatorial barcodes, except that they have spatial primers instead of target-specific regions. In some embodiments, the number of the second plurality of combinatorial barcodes is, for example, 1 / 10, 1 / 100, or 1 / 1,000 of the number of the first plurality of combinatorial barcodes. In some embodiments, the second plurality of combinatorial barcodes are used to introduce spatial labels via universal PCR. A second set of combinatorial barcodes may be added by printing or dispensing them into each well. In some embodiments, the second set of combinatorial barcodes may be delivered by the second bead into the same well containing the first bead.

[0075] Component barcode As disclosed herein, two or more component barcodes, each containing a barcode subunit sequence, can be linked together via one or more linker sequences to generate a combinatorial barcode. For example, each component barcode may include a barcode subunit sequence and one or more linker sequences or their complements, and two or more component barcodes are configured to generate a set of combinatorial barcodes by linking together via one or two linker sequences or their complements. A component barcode may include a barcode subunit sequence, a linker sequence, a target-specific region, or any combination thereof. For example, a component barcode may include one of the following configurations: a) barcode subunit array-linker array (i.e., the component barcode includes both a barcode unit array and a linker array, with the barcode subunit array located in the 5' portion of the component barcode relative to the linker array), b) complement of the linker array-barcode subunit array-linker array, c) complement of the linker array-barcode subunit array, d) barcode subunit array-linker array complement, e) linker array-barcode subunit array-linker array complement, or f) linker array-barcode subunit array. In some embodiments, the barcode subunit array and linker array can be bonded not directly by chemical bonds. For example, the barcode subunit array and linker array can be bonded via chemical or biological parts.

[0076] Component barcode reagents or combinatorial barcode reagents may include a subunit code section containing a barcode subunit sequence. A combinatorial barcode reagent may include one, two, three, four, or five or more subunit code sections. A combinatorial barcode reagent may include at least one, two, three, four, or five subunit code sections. In some cases, a combinatorial barcode reagent has one subunit code section. A subunit code section may contain a subunit code sequence. The terms subunit code sequence or barcode subunit sequence, used synonymously, may mean a unique sequence of nucleotides in a combinatorial barcode reagent. A subunit code section may be 5, 10, 15, 20, 25, 30, 35, 40, or 45 nucleotides long, or more, or at least such a length. A subunit code section may be at most 5, 10, 15, 20, 25, 30, 35, 40, or 45 nucleotides long. In some cases, the subunit coding section is 6 nucleotides long.

[0077] In some embodiments, each position in a subunit code section has a choice of four nucleotides (e.g., A, T, C, and G). If each subunit code section has an n-nucleotide length (where n is a positive integer), the number of possible unique subunit code sequences in a subunit code section is 4. n For example, if a subunit coding section is 6 nucleotides long and each position has four nucleotide options, the total number of possible unique subunit coding sequences can be 4,096 codes.

[0078] A subset of possible subunit coding sequences can be used as combinatorial barcode reagents of this disclosure. The subset of possible subunit coding sequences can be selected based on their error-correcting properties. For example, a subset of possible subunit coding sequences may have sequences that are sufficiently distinct (e.g., sufficiently different, with a significant distance) that the subunit coding sequences can correct base errors (e.g., one or two base errors). Another example is that a subset of subunit coding sequences may be of length and number in which the error-correcting barcodes of this disclosure (e.g., Hamming codes) are incorporated into the sequence.

[0079] The number of subunit coding sequences in a subset can be 5, 10, 15, 20, 25, 30, 35, 40, or 45 or more, or at least such a number. In some embodiments, the number of subunit coding sequences in a subset is at most 5, 10, 15, 20, 25, 30, 35, 40, or 45. In some embodiments, the number of subunit coding sequences in a subset is 24. Subunit coding sequences can be different from one another. Subunit coding sequences can differ by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides or more, or at least such nucleotides. In some embodiments, subunit coding sequences can differ by at most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. Subunit coding sequences can not cross-hybridize with sets of sequences. Subunit coding sequences can be error-correcting.

[0080] A combinatorial barcode reagent may contain one or more linkers. A linker can be used to integrally link combinatorial barcode reagents (e.g., to integrally concatenate combinatorial barcode reagents). A linker can be configured to hybridize to a linker of another combinatorial barcode reagent (e.g., to concatenate combinatorial barcode reagents). A linker may be on the 3' end of the subunit code of a combinatorial barcode reagent. A linker may be on the 5' end of the subunit code of a combinatorial barcode reagent. A linker may be on both the 3' and 5' ends of the subunit code of a combinatorial barcode reagent. A linker may be 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides long or more, or at least such a nucleotide length. In some embodiments, the linker may be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides long or more. In some embodiments, the linker sequence of the first combinatorial barcode reagent is hybridizable to the linker sequence of the second combinatorial barcode reagent. In some embodiments, the linker sequence of the first combinatorial barcode reagent may be a complement to the linker sequence of the second combinatorial barcode reagent. In some embodiments, the linker sequence of the first combinatorial barcode reagent is hybridizable to the linker sequence of the second combinatorial barcode reagent having 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 mismatches or more, or at least such mismatches. In some embodiments, the linker sequence of the first combinatorial barcode reagent is hybridizable to the linker sequence of the second combinatorial barcode reagent, having at most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more mismatches.

[0081] In some embodiments, the combinatorial barcode reagent includes a target-specific region. The target-specific region can be associated with (e.g., hybridized with) a target polynucleotide of the Disclosure (e.g., a single-cell polynucleotide). The target-specific region may include an oligo-dT, a gene-specific sequence, or a random multimer sequence. The target-specific region can hybridize to the sense strand of a double-stranded target polynucleotide of the Disclosure. The target-specific region can hybridize to the antisense strand of a double-stranded target polynucleotide of the Disclosure. The target-specific region may be 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides long or more, or at least such a nucleotide length. In some embodiments, the target-specific region may be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides long or more.

[0082] In some embodiments, the combinatorial barcode reagent includes one or more cell labels, molecular labels, universal labels, target-binding regions, or any labels or any combination thereof as described herein.

[0083] Combinatorial barcoding reagents can be composed of any type of nucleic acid (e.g., PNA, LNA). Combinatorial barcoding reagents can be bound to solids or semi-carriers (e.g., beads, gel particles, antibodies, hydrogels, agarose). Combinatorial barcoding reagents can be immobilized on the substrates of this disclosure (e.g., arrays). Combinatorial barcoding reagents can be incorporated into biological packages such as viruses, liposomes, and microspheres. Combinatorial reagents may contain certain parts. Some parts can function as identifiers (e.g., fluorescent parts, radioactive parts).

[0084] Set of component barcodes Some embodiments disclosed herein provide compositions comprising a set of component barcodes for generating a set of combinatorial barcodes. The number of unique combinatorial barcode reagents in a set can be n × m, where n is the number of subunit code sequences (n) and m is the number of subunit code sections (m) in the combinatorial barcode, and n and m are positive integers. For example, if there are 24 subsets of subunit code sequences and there are 4 subunit code sections in the combinatorial barcode, a total of 96 unique combinatorial barcode reagents are possible. The maximum number of sets of unique combinatorial barcodes that can be generated using n × m unique component barcodes is nm, where n is the number of subunit code sequences and m is the number of subunit code sections in the combinatorial barcode. For example, if a combinatorial barcode has 24 subunit code sequences and 4 subunit code sections, a maximum set of 331,776 unique combinatorial barcodes can be generated from a set of 96 unique combinatorial barcode reagents.

[0085] The number of subunit code sections in a combinatorial barcode can be 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more, or at least such a number. The number of subunit code sections in a combinatorial barcode can be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In some embodiments, the number of subunit code sections in a combinatorial barcode is 4 (as shown in Figure 1).

[0086] Some combinatorial barcode reagents in a set of combinatorial barcode reagents may contain target-specific regions. The number of combinatorial barcode reagents in a set containing target-specific regions may be 1, 2, 3, 4, or 5 or more, or at least such a number. In some embodiments, the number of combinatorial barcode reagents in a set containing target-specific regions may be at most 1, 2, 3, 4, or 5. As shown in Figures 2A-C, in some embodiments, the number of combinatorial barcode reagents in a set containing target-specific regions is at least 2. In some cases, the number of combinatorial barcode reagents in a set containing target-specific regions is 2 (Figures 2A-B). In some embodiments, the number of combinatorial barcode reagents in a set containing target-specific regions is at least 1. In some cases, the number of combinatorial barcode reagents in a set containing target-specific regions is 1 (Figure 2C).

[0087] A set of combinatorial barcode reagents may include several combinatorial barcode reagents containing target-specific regions and several combinatorial barcode reagents not containing target-specific regions. In some cases, a set of combinatorial barcode reagents may include one combinatorial barcode reagent having a target-specific region, while the remaining combinatorial barcode reagents may not have a target-specific region (for example, they may include subunit code sections / sequences and one or more linkers). In some cases, a set of combinatorial barcode reagents may include two combinatorial barcode reagents having target-specific regions, while the remaining combinatorial barcode reagents may not have a target-specific region (for example, they may include subunit code sections / sequences and one or more linkers).

[0088] The number of combinatorial barcode reagents in a set, which may not have a target-specific region, can be 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more, or at least such a number. In some embodiments, the number of combinatorial barcode reagents in a set, which may not have a target-specific region, can be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10.

[0089] A set of combinatorial barcode reagents can encode multiple combinatorial barcodes. A combinatorial barcode may contain the entire sequence of the combinatorial barcode reagents in the set (i.e., for a given target polynucleotide). A combinatorial barcode may be the entire sequence of a concatenate combinatorial barcode reagent for a given target polynucleotide (e.g., it may be bound and integrated via a linker). A combinatorial barcode may contain the linker sequence of the combinatorial barcode reagent. A combinatorial barcode may exclude the linker sequence of the combinatorial barcode reagent (e.g., it may contain only the subunit coding sequence). A combinatorial barcode may be the entire sequence of the subunit coding section of the combinatorial barcode reagent for a given target polynucleotide.

[0090] In some embodiments, a combinatorial barcode may mean (for example, can be formed using) two or more combinatorial barcode reagents that are hybridized to one or more linker sequences and / or hybridized to a target polynucleotide via one or more target-specific regions. In some embodiments, a combinatorial barcode may mean two or more combinatorial barcode reagents incorporated into a single polynucleotide with or without the target polynucleotide via extension, reverse transcription, and / or amplification.

[0091] In some embodiments, the combinatorial barcode may be bipertite. A portion of the combinatorial barcode is ligable to the 5' end of the target polynucleotide. A portion of the combinatorial barcode is ligable to the 3' end of the target polynucleotide. The entire combinatorial barcode is ligable to the 5' end of the target polynucleotide. The entire combinatorial barcode is ligable to the 3' end of the target polynucleotide.

[0092] Combinatorial barcodes can be constructed by combining combinatorial barcode reagents. A combinatorial barcode may include one or more cell labels, molecular labels, universal labels, target-binding regions, any labels disclosed herein, or any combination thereof. In some cases, combinatorial barcode reagents may not include labels. For example, if a combinatorial barcode consists of four combinatorial barcode reagents, one of the reagents may not include a label, and three of the reagents may include any labels disclosed herein. Combinatorial barcode reagents may include probabilistic barcodes disclosed herein.

[0093] The set of combinatorial barcode reagents can be random. For example, all or part of the linker sequences in a combinatorial barcode may be identical. Therefore, component barcodes can be joined in any order. The set of combinatorial barcode reagents can also be non-random. For example, all the linker sequences in a combinatorial barcode may be different. Therefore, the order of component barcodes can be joined in a specific order.

[0094] In some embodiments, the combinatorial barcode reagent is bound to any solid or semi-solid carrier disclosed herein. For example, the combinatorial barcode reagent can be bound to a combination of beads and gel particles, or to a combination of a second combinatorial barcode reagent bound to beads and a first combinatorial barcode reagent in solution, or to a combination of a first combinatorial barcode reagent immobilized on a substrate in a microwell and a second combinatorial barcode reagent in solution in a microwell or present on solid particles. In some embodiments, the combinatorial barcode reagent is embedded in a hydrogel or similar material capable of acting as a sponge-like scaffold for the purpose of localizing biological samples and nucleic acids.

[0095] Solid carriers The combinatorial barcode reagents and / or combinatorial barcodes disclosed herein are conjugable to solid supports (e.g., beads, substrates). The combinatorial barcode reagents and / or combinatorial barcodes disclosed herein may be located on solid supports (e.g., in microwells of an array). As used herein, the terms “tethered,” “conjugated,” and “immobilized” are used synonymously to mean covalent or noncovalent means for conjugating a probabilistic barcode to a solid support. Any of the various different solid supports may be used as solid supports for conjugating pre-synthesized combinatorial barcode reagents or for in situ solid-phase synthesis of combinatorial barcode reagents.

[0096] In some cases, the solid carrier is a bead. The bead may encompass any type of solid, porous, or hollow sphere, ball, bearing, cylinder, or other similar construct made of plastic, ceramic, metal, or polymer material capable of immobilizing nucleic acids (e.g., covalently or non-covalently). The bead may include discrete particles that are spherical (e.g., microspheres) or non-spherical or irregular in shape, such as cubic, rectangular, pyramidal, cylindrical, conical, oblate, or disk-shaped. Beads can have a non-spherical shape.

[0097] The beads may include, but are not limited to, a variety of materials such as paramagnetic materials (e.g., magnesium, molybdenum, lithium, and tantalum), superparamagnetic materials (e.g., ferrite (Fe3O4, magnetite) nanoparticles), ferromagnetic materials (e.g., iron, nickel, cobalt, some alloys thereof, and some rare earth metal compounds), ceramics, plastics, glass, polystyrene, silica, methylstyrene, acrylic polymers, titanium, latex, Sepharose, agarose, hydrogels, polymers, cellulose, nylon, and any combination thereof.

[0098] The diameter of the beads may be 5 μm, 10 μm, 20 μm, 25 μm, 30 μm, 35 μm, 40 μm, 45 μm, or 50 μm, or roughly at least such a length. The diameter of the beads may be at most about 5 μm, 10 μm, 20 μm, 25 μm, 30 μm, 35 μm, 40 μm, 45 μm, or 50 μm. The diameter of the beads can be correlated with the diameter of the wells in the substrate. For example, the diameter of the beads may be 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% longer or shorter than the diameter of the well, or at least such a length. In some embodiments, the diameter of the beads may be at most 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% longer or shorter than the diameter of the well. The diameter of the beads can be related to the diameter of the cells (for example, single cells trapped in wells of the substrate). The diameter of the beads may be 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, or 300% or more longer or shorter than the diameter of the cells, or at least such a length. In some embodiments, the diameter of the beads may be at most 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, or 300% or more longer or shorter than the diameter of the cells.

[0099] The beads can be embedded in and / or bonded to the substrates of this disclosure. The beads can be embedded in and / or bonded to gels, hydrogels, polymers, and / or matrices. The spatial position of the beads within the substrate (e.g., gel, matrix, scaffold, or polymer) can be identified using a spatial label present in a probabilistic barcode on the bead that can function as a positional address.

[0100] Examples of beads include, but are not limited to, streptavidin beads, agarose beads, magnetic beads, Dynabead®, MACS® microbeads, antibody conjugate beads (e.g., anti-immunoglobulin microbeads), protein A conjugate beads, protein G conjugate beads, protein A / G conjugate beads, protein L conjugate beads, oligo dT conjugate beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, and BcMag® carboxy-terminated magnetic beads.

[0101] The beads can be associated with quantum dots or fluorescent dyes (e.g., impregnation with them) to fluoresce in one or more optical channels. The beads can be associated with iron oxide or chromium oxide to be paramagnetic or ferromagnetic. The beads can be identifiable. The beads can be imaged using a camera. The beads can have a detectable code associated with them. For example, the beads can contain RFID tags. The beads can contain any detectable tag (e.g., UPC code, electronic barcode, etching identifier). The beads can change size due to swelling in organic or inorganic solutions, for example. The beads can be hydrophobic or hydrophilic. The beads can be biocompatible.

[0102] The solid support (e.g., beads) is visible. The solid support may contain a visibility tag (e.g., a fluorescent dye). The solid support (e.g., beads) is etchable with an identifier (e.g., a number). The identifier is visible by imaging the solid support (e.g., beads).

[0103] Solid supports may contain insoluble, semi-soluble, or insoluble materials. A solid support can also be referred to as "functionalized" if it contains a linker, scaffold, building block, or other reactive moiety bonded to it, while a solid support can be "non-functionalized" if it lacks such reactive moiety bonded to it. Solid supports are available in a free state in solution, for example, in a microtiter well system, a flow-through system such as a column, or a dipstick setting.

[0104] Solid carriers may include membranes, paper, plastics, coated surfaces, flat surfaces, glass, slides, chips, or any combination thereof. Solid carriers may take the form of resins, gels, microspheres, or other geometric constructs. Solid carriers may include silica chips, microparticles, nanoparticles, plates, arrays, capillaries, flat carriers, e.g., glass fiber filters, glass surfaces, metal surfaces (steel, gold, silver, aluminum, silicon, and copper), glass carriers, plastic carriers, silicon carriers, chips, filters, membranes, microwell plates, slides, multiwell plates or membranes (e.g., formed from polyethylene, polypropylene, polyamide, polyvinylidene difluoride) and / or wafers, combs, wafers (e.g., silicon wafers) with flat surfaces, pits or arrays of nanoliter wells, pins or needles (e.g., arrays of pins suitable for combinatorial synthesis or analysis) or beads, wafers with or without filter bottoms, and wafers with pits.

[0105] The solid carrier may include a polymer matrix (e.g., a gel, a hydrogel). The polymer matrix may be permeable to intracellular spaces (e.g., around organelles). The polymer matrix may be pumpable throughout the circulatory system.

[0106] Solid carriers can be biological molecules. For example, solid carriers can be nucleic acids, proteins, antibodies, histones, cell compartments, lipids, carbohydrates, etc. Solid carriers that are biological molecules can be amplified, translated, transcribed, degraded, and / or modified (e.g., pegylated, SUMOlated, acetylated, methylated). Solid carriers that are biological molecules can provide spatial and temporal information in addition to the spatial label bound to the biological molecule. For example, a biological molecule may have a first conformation when unmodified, but may change to a second conformation when modified. The probabilistic barcodes of this disclosure can be exposed to a target in various conformations. For example, a biological molecule may have a probabilistic barcode that becomes inaccessible due to the folding of the biological molecule. When a biological molecule is modified (e.g., acetylated), the biological molecule can change its conformation to expose the probabilistic label. The timing of the modification can provide another temporal dimension to the probabilistic barcoding method of this disclosure.

[0107] In some embodiments, the biological molecule containing the combinatorial barcoding reagent of this disclosure may be located in the cytoplasm of a cell. Upon activation, the biological molecule is capable of moving to the nucleus and subsequently performing probabilistic barcoding. Thus, modification of the biological molecule allows for the encoding of additional spatial-temporal information for targets identified by probabilistic barcoding.

[0108] Dimensional labels can provide spatial-temporal information about biological events (e.g., cell division). For example, a dimensional label can be added to a first cell, which can then divide to produce a second daughter cell, which may contain all or part of the dimensional label, or none at all. The dimensional label can be activated in both the progenitor and daughter cells. Thus, the dimensional label can provide temporal information about combinatorial barcoding in a discernible space.

[0109] Base material The substrate may mean a certain type of solid carrier. The substrate may mean a solid carrier that may contain the combinatorial barcode reagent of this disclosure. The substrate may comprise a plurality of microwells. The microwells may comprise small reaction chambers of a specified volume. The microwells may contain one or more cells. The microwells may contain only one cell. The microwells may contain one or more solid carriers. The microwells may contain only one solid carrier. In some cases, the microwells contain a single cell and a single solid carrier (e.g., a bead). The microwells may contain the combinatorial barcode reagent of this disclosure.

[0110] Microwells in an array can be fabricated in a variety of shapes and sizes. The geometry of the wells is not limited to, but may include cylinders, cones, hemispheres, rectangles, or polyhedra (e.g., three-dimensional geometries composed of several planes, such as hexagonal prisms, octagonal prisms, inverted triangular pyramids, inverted square pyramids, inverted pentagonal pyramids, inverted hexagonal pyramids, or inverted truncated pyramids). Microwells may include shapes that combine two or more of these geometries. For example, a microwell may be partially cylindrical with the remaining part being an inverted cone. A microwell may contain two side-by-side cylinders, one with a diameter longer than the other (e.g., roughly corresponding to the diameter of a cell) (e.g., roughly corresponding to the diameter of a bead), and they are connected by a vertical channel (i.e., parallel to the cylinder axis) that runs along the entire length (depth) of the cylinders. The opening of a microwell may be located on the upper surface of the substrate. The opening of a microwell may be located on the lower surface of the substrate. The closed end (or bottom) of a microwell may be flat. The closed end (or bottom) of a microwell may have a curved surface (e.g., convex or concave). The shape and / or size of the microwell may be determined based on the type of cells or solid carrier to be trapped within the microwell.

[0111] The substrate portion between wells may have a certain topology. For example, the substrate portion between wells may be round. The substrate portion between wells may be pointed. The space portion of the substrate between wells may be flat. The substrate portion between wells does not have to be flat. In some cases, the substrate portion between wells is round. In other words, the substrate portion that does not contain wells may have a curved surface. The curved surface can be fabricated such that the highest point of the curved surface (e.g., a vertex) is the furthest point between the edges of two or more wells (e.g., equidistant from the wells). The curved surface can be fabricated such that the starting point of the curved surface is on the edge of a first microwell and ends at the edge of a second microwell, forming a parabola. This parabola can extend two-dimensionally to capture adjacent microwells on a hexagonal grid of wells. The curved surface can be fabricated such that the surface between wells is higher and / or curved than the plane of the well openings. The height of the curved surface can be 0.1, 0.5, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, or 7 micrometers or more, or at least such a height. In some embodiments, the height of the curved surface can be at most 0.1, 0.5, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, or 7 micrometers or more.

[0112] The dimensions of a microwell can be characterized by its diameter and depth. As used herein, the diameter of a microwell means the largest circle that can be inscribed in the planar cross-section of the microwell's geometry. The diameter of a microwell can range from about 1 to about 10 times the diameter of the cell or solid carrier trapped in the microwell. The diameter of a microwell can be 1 or at least 1, at least 1.5, at least 2, at least 3, at least 4, at least 5, or at least 10 times the diameter of the cell or solid carrier trapped in the microwell. In some embodiments, the diameter of a microwell can be as much as 10, as much as 5, as much as 4, as much as 3, as much as 2, as much as 1.5, or as much as 1 time the diameter of the cell or solid carrier trapped in the microwell. The diameter of a microwell can be about 2.5 times the diameter of the cell or solid carrier trapped in the microwell.

[0113] The diameter of a microwell can be defined by absolute dimensions. The diameter of a microwell can be in the range of approximately 5 to approximately 60 micrometers. The diameter of a microwell can be 5 micrometers or at least 5 micrometers, at least 10 micrometers, at least 15 micrometers, at least 20 micrometers, at least 25 micrometers, at least 30 micrometers, at least 35 micrometers, at least 40 micrometers, at least 45 micrometers, at least 50 micrometers, or at least 60 micrometers. The diameter of a microwell can be at most 60 micrometers, at most 50 micrometers, at most 45 micrometers, at most 40 micrometers, at most 35 micrometers, at most 30 micrometers, at most 25 micrometers, at most 20 micrometers, at most 15 micrometers, at most 10 micrometers, or at most 5 micrometers. The diameter of a microwell can be approximately 30 micrometers.

[0114] The depth of the microwells may be selected to provide efficient trapping of cells and solid carriers. The depth of the microwells may be selected to provide efficient exchange of assay buffers and other reagents contained within the wells. The diameter-to-height ratio (i.e., aspect ratio) may be selected so that when cells and solid carriers are placed in the microwells, they do not displace beyond the microwells due to fluid motion. The dimensions of the microwells may be selected so that the microwells have sufficient space to accommodate solid carriers and cells of various sizes without being removed beyond the microwells due to fluid motion. The depth of the microwells may be in the range of approximately 1 to 10 times the diameter of the cells or solid carriers trapped in the microwells. The depth of the microwells may be 1 or at least 1, at least 1.5, at least 2, at least 3, at least 4, at least 5, or at least 10 times the diameter of the cells or solid carriers trapped in the microwells. The depth of the microwell can be at most 10 times, at most 5 times, at most 4 times, at most 3 times, at most 2 times, at most 1.5 times, or at most 1 time, the diameter of the cells or solid carrier trapped in the microwell. The depth of the microwell can be approximately 2.5 times the diameter of the cells or solid carrier trapped in the microwell.

[0115] The depth of a microwell can be defined by its absolute dimensions. The depth of a microwell can be in the range of approximately 10 to approximately 60 micrometers. The depth of a microwell can be 10 micrometers or at least 10 micrometers, at least 20 micrometers, at least 25 micrometers, at least 30 micrometers, at least 35 micrometers, at least 40 micrometers, at least 50 micrometers, or at least 60 micrometers. The depth of a microwell can be at most 60 micrometers, at most 50 micrometers, at most 40 micrometers, at most 35 micrometers, at most 30 micrometers, at most 25 micrometers, at most 20 micrometers, or at most 10 micrometers. The depth of a microwell can be approximately 30 micrometers.

[0116] The volume of the microwells used in the methods, devices, and systems of this disclosure is approximately 200 micrometers. 3 ~Approximately 120,000 micrometers 3 It may be within this range. The volume of a microwell is at least 200 micrometers. 3 at least 500 micrometers 3 at least 1,000 micrometers 3 at least 10,000 micrometers 3 at least 25,000 micrometers 3 at least 50,000 micrometers 3 at least 100,000 micrometers 3 , or at least 120,000 micrometers 3 This is possible. The volume of a microwell is at most 120,000 micrometers. 3 at most 100,000 micrometers 3 at most 50,000 micrometers 3 at most 25,000 micrometers 3 at most 10,000 micrometers 3at most 1,000 micrometers 3 at most 500 micrometers 3 , or at most 200 micrometers 3 This is possible. The volume of a microwell is approximately 25,000 micrometers. 3 The volume of a microwell can be any range bounded by any of these values ​​(for example, about 18,000 micrometers). 3 ~Approximately 30,000 micrometers 3 ) may be included in.

[0117] The microwell volumes are 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nanoliters. 3 Or it could be more or at least that volume. The volume of a microwell is at most 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nanoliters. 3 Or it may be more. The volume of liquid that can be contained in a microwell is at least 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nanoliters. 3 Or it may be even more. The volume of liquid that can be contained in a microwell is at most 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nanoliters. 3 Or it can be more. Microwell volumes are 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 picoliters. 3 Or it can be more or at least that volume. Microwell volumes are at most 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 picoliters. 3 Or it may be more. The volume of liquid that can be contained in a microwell is at least 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 picoliters. 3 Or it may be more. The volume of liquid that can be contained in a microwell is at most 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 picoliters. 3 Or it could be even more.

[0118] The volumes of the microwells used in the methods, devices, and systems of this disclosure may be further characterized by the volume variation between the microwells. The coefficient of variation (expressed as a percentage) of the microwell volumes may be in the range of about 1% to about 10%. The coefficient of variation of the microwell volumes may be at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, or at least 10%. The coefficient of variation of the microwell volumes may be at most 10%, at most 9%, at most 8%, at most 7%, at most 6%, at most 5%, at most 4%, at most 3%, at most 2%, or at most 1%. For example, the coefficient of variation of the microwell volumes may have any value within the range enclosed by these values, e.g., about 1.5% to about 6.5%. In some embodiments, the coefficient of variation of the microwell volumes may be about 2.5%.

[0119] The ratio of the volume of the microwells used in the methods, devices, and systems of this disclosure to the surface area of ​​the beads (or the surface area of ​​the solid support to which the probabilistic barcode oligonucleotides can be bound) may be in the range of about 2.5 to about 1,520 micrometers. The ratio may be at least 2.5, at least 5, at least 10, at least 100, at least 500, at least 750, at least 1,000, or at least 1,520. The ratio may be at most 1,520, at most 1,000, at most 750, at most 500, at most 100, at most 10, at most 5, or at most 2.5. The ratio may be about 67.5. The ratio of the volume of the microwells to the surface area of ​​the beads (or the solid support used for fixation) may fall within any range bounded by any of these values ​​(e.g., about 30 to about 120).

[0120] The wells of a microwell array can be arranged in a one-dimensional, two-dimensional, or three-dimensional array. In some embodiments, a three-dimensional array can be achieved, for example, by stacking a series of two or more two-dimensional arrays (i.e., by stacking two or more substrates containing microwell arrays).

[0121] The pattern and spacing between microwells can be selected to optimize the efficiency of trapping single cells and single solid carriers (e.g., beads) in each well, and to maximize the number of wells per unit area of ​​the array. Microwells can be distributed according to various random or non-random patterns. For example, they can be distributed randomly across the entire surface of the array substrate, or arranged in a square grid, rectangular grid, hexagonal grid, etc. In some cases, microwells can be arranged in a hexagonal pattern. The center-to-center distance (or space) between wells can vary from about 5 micrometers to about 75 micrometers. In some cases, the space between microwells is about 10 micrometers. In other embodiments, the space between wells is at least 5 micrometers, at least 10 micrometers, at least 15 micrometers, at least 20 micrometers, at least 25 micrometers, at least 30 micrometers, at least 35 micrometers, at least 40 micrometers, at least 45 micrometers, at least 50 micrometers, at least 55 micrometers, at least 60 micrometers, at least 65 micrometers, at least 70 micrometers, or at least 75 micrometers. The space between microwells can be at most 75 micrometers, at most 70 micrometers, at most 65 micrometers, at most 60 micrometers, at most 55 micrometers, at most 50 micrometers, at most 45 micrometers, at most 40 micrometers, at most 35 micrometers, at most 30 micrometers, at most 25 micrometers, at most 20 micrometers, at most 15 micrometers, at most 10 micrometers, or at most 5 micrometers. The space between microwells can be about 55 micrometers. The space of a microwell can fall within any range bounded by any of these values ​​(for example, from about 18 micrometers to about 72 micrometers).

[0122] Microwell arrays may include inter-microwell surface features designed to help guide cells and solid carriers into the wells and / or prevent them from settling on the inter-well surfaces. Examples of preferred surface features, but not limited to, include dome-shaped, ridge-shaped, or peak-shaped surface features surrounding the wells or spanning the inter-well surfaces.

[0123] The total number of wells in a microwell array can be determined by the well pattern and spacing, as well as the overall dimensions of the array. The number of microwells in an array can be in the range of approximately 96 to approximately 5,000,000 or more. The number of microwells in an array can be at least 96, at least 384, at least 1,536, at least 5,000, at least 10,000, at least 25,000, at least 50,000, at least 75,000, at least 100,000, at least 500,000, at least 1,000,000, or at least 5,000,000. The number of microwells in an array can be at most 5,000,000, at most 1,000,000, at most 75,000, at most 50,000, at most 25,000, at most 10,000, at most 5,000, at most 1,536, at most 384, or at most 96 wells. The number of microwells in an array can be approximately 96, 384, and / or 1536. The number of microwells can be approximately 150,000. The number of microwells in an array can fall within any range bounded by any of these values ​​(e.g., approximately 100 to 325,000).

[0124] Microwell arrays can be fabricated using one of several fabrication techniques. Examples of possible fabrication methods include, but are not limited to, bulk micromachining techniques such as photolithography and wet chemical etching, plasma etching, or deep reactive ion etching, micromolding and microembossing, laser micromachining, 3D printing, or other direct writing fabrication processes using curable materials, and similar techniques.

[0125] Microwell arrays can be fabricated from one of several substrate materials. The choice of material may depend on the choice of fabrication technique, and vice versa. Examples of suitable materials, but not limited to, include silicon (fused silica), glass, polymers (e.g., agarose, gelatin, hydrogel, polydimethylsiloxane (PDMS, elastomer), polymethyl methacrylate (PMMA), polycarbonate (PC), polypropylene (PP), polyethylene (PE), high-density polyethylene (HDPE), polyimide, cyclic olefin polymer (COP), cyclic olefin copolymer (COC), polyethylene terephthalate (PET), epoxy resins, thiol-ene resins, and metals or metal films (e.g., aluminum, stainless steel, copper, nickel, chromium, and titanium). In some cases, the microwells contain optical adhesives. In some cases, the microwells are fabricated with optical adhesives. In some cases, Microwell arrays are fabricated with and / or PDMS. In some cases, microwells are fabricated with plastic. Hydrophilic materials may be desirable for the fabrication of microwell arrays (e.g., to improve wettability and to minimize nonspecific binding of cells and other biomaterials). Hydrophobic materials that can be treated or coated (e.g., by oxygen plasma treatment or grafting of polyethylene oxide surface layers) are also available. The use of porous hydrophilic materials for the fabrication of microwell arrays may be desirable to facilitate capillary wicking / venting of trapped bubbles within the device. Microwell arrays can be fabricated from a single material. Microwell arrays may contain two or more different materials that are bonded together or mechanically linked.

[0126] Microwell arrays can be fabricated using substrates of various sizes and shapes. For example, the shape (or footprint) of the substrate on which the microwells are fabricated can be square, rectangular, circular, or irregular. The footprint of the microwell array substrate can be similar to that of a microtiter plate. The footprint of the microwell array substrate can be similar to that of a standard microscope slide. For example, it can be approximately 75 mm long × 25 mm wide (approximately 3 inches long × 1 inch wide), or approximately 75 mm long × 50 mm wide (approximately 3 inches long × 2 inches wide). The thickness of the substrate on which the microwells are fabricated can be in the range of approximately 0.1 mm to approximately 10 mm or greater. The thickness of the microwell array substrate can be at least 0.1 mm thick, at least 0.5 mm thick, at least 1 mm thick, at least 2 mm thick, at least 3 mm thick, at least 4 mm thick, at least 5 mm thick, at least 6 mm thick, at least 7 mm thick, at least 8 mm thick, at least 9 mm thick, or at least 10 mm thick. The thickness of the microwell array substrate can be at most 10 mm, at most 9 mm, at most 8 mm, at most 7 mm, at most 6 mm, at most 5 mm, at most 4 mm, at most 3 mm, at most 2 mm, at most 1 mm, at most 0.5 mm, or at most 0.1 mm. The thickness of the microwell array substrate can be approximately 1 mm. For example, the thickness of the microwell array substrate can be any value within these ranges. For example, the thickness of the microwell array substrate can be approximately 0.2 mm to approximately 9.5 mm. The thickness of the microwell array substrate can be uniform.

[0127] Various surface treatment and surface modification techniques can be used to alter the properties of the microwell array surface. Examples, but not limited to, include oxygen plasma treatment to make hydrophobic material surfaces more hydrophilic; the use of wet or dry etching techniques to smooth (or roughen) glass and silicon surfaces; the adsorption or grafting of polyethylene oxide or other polymer layers (e.g., Pluronic) or bovine serum albumin onto the substrate surface to increase hydrophilicity and reduce nonspecific adsorption of biomolecules and cells; and the use of silane reactions to graft chemically reactive functional groups onto silicon and glass surfaces to otherwise deactivate them. Photodeprotection techniques can be used to selectively activate chemically reactive functional groups at specific locations on the array structure; for example, selective addition or activation of chemically reactive functional groups such as primary amine groups or carboxyl groups on the inner walls of microwells can be used to covalently bond oligonucleotide probes, peptides, proteins, or other biomolecules to the microwell walls. The choice of surface treatment or surface modification used may depend on both or both of the desired surface properties and the type of material used to fabricate the microwell array.

[0128] To prevent cross-hybridization of target nucleic acids between adjacent microwells, the openings of the microwells may be sealed, for example, during the cell lysis process. For example, the microwells (or arrays of microwells) may be sealed or capped using a flexible membrane or sheet (i.e., plate or platen) of solid material that clamps the surface of the microwell array substrate, or using suitable beads if the diameter of the beads is larger than the diameter of the microwells.

[0129] Seals formed using flexible membranes or sheets of solid materials may include, for example, inorganic nanoporous membranes (e.g., aluminum oxide), dialysis membranes, glass slides, coverslips, elastomer films (e.g., PDMS), or hydrophilic polymer films (e.g., polymer films coated with a thin film of agarose hydrated in a solubility buffer).

[0130] The solid carrier (e.g., beads) used for capping the microwells may include any of the solid carriers of this disclosure (e.g., beads). In some cases, the solid carrier is a crosslinked dextran bead (e.g., Sephadex). The crosslinked dextran may be in the range of about 10 micrometers to about 80 micrometers. The crosslinked dextran beads used for capping may be in the range of 20 micrometers to about 50 micrometers. In some embodiments, the beads may be at least about 10, 20, 30, 40, 50, 60, 70, 80, or 90% larger than the diameter of the microwell. The beads used for capping may be at most about 10, 20, 30, 40, 50, 60, 70, 80, or 90% larger than the diameter of the microwell.

[0131] Seals or caps can prevent the migration of macromolecules (e.g., nucleic acids) from the wells while allowing buffer to enter and exit the microwells. At least approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides or more of macromolecules can be prevented from entering or leaving the microwells by seals or caps. At most approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides or more of macromolecules can be prevented from entering or leaving the microwells by seals or caps.

[0132] Solid carriers (e.g., beads) can be distributed into a substrate. The solid carriers (e.g., beads) can be distributed into the wells of the substrate, removed from the wells of the substrate, or transported via a device containing one or more microwell arrays by centrifugation or other non-magnetic means. Microwells of the substrate can be preloaded with solid carriers. Microwells of the substrate can hold at least one, two, three, four, or five or more solid carriers. Microwells of the substrate can hold at most one, two, three, four, or five or more solid carriers. In some cases, microwells of the substrate can hold one solid carrier.

[0133] Individual cells and beads can be compartmentalized in microwells using alternative methods. For example, a single solid carrier and a single cell can be confined within a single droplet of emulsion (e.g., in droplet digital microfluidic systems).

[0134] The cells can potentially be confined within porous beads that themselves contain multiple tethered probabilistic barcodes. Individual cells and solid carriers can be compartmentalized into any type of container, microcontainer, reaction chamber, reaction vessel, etc.

[0135] Single-cell combinatorial barcoding can be performed without the use of microwells. Single-cell combinatorial barcoding assays can be performed without any physical containers. For example, containerless combinatorial barcoding can be performed by embedding cells and beads in close proximity to each other within a polymer or gel layer to form a diffusion barrier between different cell / bead pairs. Another example is that containerless combinatorial barcoding can be performed in situ, in vivo, on intact solid tissue, on intact cells, and / or subcellularly.

[0136] Microwell arrays can be consumer components of an assay system. Microwell arrays can be reusable. Microwell arrays can be configured to be used as standalone devices for manually performing assays, or they can be configured to include fixed or removable components of an instrumentation system that provides full or partial automation of the assay procedure. In some embodiments of the methods of this disclosure, a bead-based library of probabilistic barcodes can be deposited in the wells of a microwell array as part of the assay procedure. In some embodiments, the beads can be pre-loaded in the wells of a microwell array and can also be provided to the user as part of a kit for performing, for example, probabilistic barcoding and digital counting of nucleic acid targets.

[0137] In some embodiments, two pairs of microwell arrays are provided, one preloaded with beads held in place by a first magnet, and the other used by the user when loading individual cells. After distributing the cells into the second microwell array, the two arrays are positioned facing each other, and the first magnet is removed while the second magnet is used to pull the beads of the first array into the corresponding microwells of the second array, thereby ensuring that the beads remain stationary on the cells in the second microwell array, thereby maximizing the efficient binding of target molecules to probabilistic barcodes on the beads while minimizing diffusion loss of target molecules after cell lysis.

[0138] The microwell arrays of this disclosure may be preloaded with a solid carrier (e.g., beads). Each well of the microwell array may contain a single solid carrier. At least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% of the wells in the microwell array may be preloaded with a single solid carrier. At most 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% of the wells in the microwell array may be preloaded with a single solid carrier. The solid carriers may contain the probabilistic barcodes and / or combinatorial barcodes of this disclosure. Cell labeling of probabilistic barcodes on different solid carriers may differ. Cell labeling of probabilistic barcodes on the same solid carrier may be identical.

[0139] three dimensional base material Three-dimensional arrays can have any shape. Three-dimensional substrates can be made from any material used for the substrates of this disclosure. In some cases, the three-dimensional substrates include DNA origami. DNA origami structures incorporate DNA as building material to create nanoscale shapes. The DNA origami process may include folding one or more long "scaffold" DNA strands into a specific shape using multiple reasonably designed "staple" DNA strands. The sequence of the staple strands can be designed to hybridize to specific portions of the scaffold strand and, in the process, to give the scaffold strand a specific shape. DNA origami may include a scaffold strand and multiple reasonably designed staple strands. The scaffold strand may have any sufficiently non-repeating sequence.

[0140] The arrangement of the staple strands can be selected such that the DNA origami has at least one shape to which a probabilistic marker can be bound. In some embodiments, the DNA origami can be any shape having at least one inner surface and at least one outer surface. The inner surface can be any surface region of the DNA origami where interaction with the sample surface is sterically prevented, and the outer surface can be any surface region of the DNA origami where interaction with the sample surface is not sterically prevented. In some embodiments, the DNA origami has one or more openings (e.g., two openings) so that particles (e.g., a solid carrier) can access the inner surface of the DNA origami. For example, in a particular embodiment, the DNA origami has one or more openings so that particles smaller than 10 micrometers, 5 micrometers, 1 micrometer, 500 nm, 400 nm, 300 μm, 250 nm, 200 nm, 150 nm, 100 nm, 75 nm, 50 nm, 45 nm, or 40 nm can contact the inner surface of the DNA origami.

[0141] DNA origami can change its shape (conformation) in response to one or more specific environmental stimuli. Therefore, a region of DNA origami can be an inner surface when the DNA origami exhibits several conformations, but an outer surface when the invention exhibits other conformations. In some embodiments, DNA origami can respond to specific environmental stimuli by exhibiting a new conformation.

[0142] In some embodiments, the staple strands of the DNA origami can be selected so that the DNA origami is substantially barrel-shaped or tubular. The staples of the DNA origami can be selected so that both ends of the barrel shape are closed, or so that one or both ends are open, allowing particles to enter the inside of the barrel and access its inner surface. In certain embodiments, the barrel shape of the DNA origami may be a hexagonal tube.

[0143] In some embodiments, the staple strands of the DNA origami are selectable such that the DNA origami has a first domain and a second domain, the first end of the first domain is bound to the first end of the second domain by one or more single-stranded DNA hinges, and the second end of the first domain is bound to the second domain of the second domain by one or more molecular latches. Multiple staples are selectable such that, when all molecular latches are in contact with their respective external stimuli, the second end of the first domain remains unglued from the second end of the second domain. A latch can be formed from two or more staples, each including at least one staple strand having at least one stimulus-binding domain capable of binding to an external stimulus such as nucleic acid, lipid, or protein, and at least one other staple strand having at least one latch domain that binds to the stimulus-binding domain. The binding of the stimulus-binding domain to the latch domain supports the stability of the first conformation of the DNA origami.

[0144] Synthesis of combinatorial barcodes on solid carriers and substrates In some embodiments, combinatorial barcode reagents can be synthesized on solid supports (e.g., beads). Pre-synthesized combinatorial barcode reagents (e.g., containing a 5' amine that can be bound to a solid support) can be bound to a solid support (e.g., beads) by any of various immobilization techniques that include functional group pairs on the solid support and on the probabilistic barcode. Combinatorial barcode reagents may contain functional groups. Solid supports (e.g., beads) may contain functional groups. Functional groups of combinatorial barcode reagents and solid supports may include, for example, biotin, streptavidin, primary amines, carboxyl, hydroxyl, aldehyde, ketone, and any combination thereof. Combinatorial barcodes can be tethered to a solid support, for example, by coupling a 5' amino group on the combinatorial barcode reagent with a carboxyl group on the functionalized solid support (e.g., using 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide). Any remaining uncoupled combinatorial barcode reagent can be removed from the reaction mixture by performing multiple rinsing steps. In some embodiments, the combinatorial barcode reagent and the solid support are indirectly linked via a linker molecule (e.g., a short functionalized hydrocarbon molecule or polyethylene oxide molecule) using similar bonding chemistry. The linker may be a cleavage linker, such as an acid-unstable linker or a photocleavage linker.

[0145] Combinatorial barcode reagents can be synthesized on solid supports (e.g., beads) using one of several solid-phase oligonucleotide synthesis techniques, such as phosphodiester synthesis, phosphotryester synthesis, phosphytotryester synthesis, and phosphoramidite synthesis. Single nucleotides can be stepwise coupled to the growing tether-linked combinatorial barcode reagent. In some embodiments, short pre-synthesized sequences (or blocks) of several oligonucleotides can be coupled to the growing tether-linked combinatorial barcode reagent.

[0146] Combinatorial barcode reagents can be synthesized by a stepwise coupling reaction or a block coupling reaction using one or more rounds of split - pool synthesis. In this case, the entire pool of synthetic beads is divided into several individual small pools, and each is then subjected to a different coupling reaction. Subsequently, the individual pools are recombined and mixed to randomize the growing combinatorial barcode reagent sequences across the entire pool of beads. Split - pool synthesis is an example of a combinatorial synthesis process that synthesizes the maximum number of chemical compounds using the minimum number of chemical coupling steps. The potential diversity of the compound library thus produced is determined by the number of unique building blocks (e.g., nucleotides) available in each coupling step and the number of coupling steps used in the production of the library. For example, a split - pool synthesis involving 10 rounds of coupling using 4 different nucleotides in each step will generate 10 = 1,048,576 unique nucleotide sequences. In some embodiments, split - pool synthesis can be performed using enzymatic methods such as polymerase extension or ligation reactions instead of chemical coupling. For example, in each round of a split - pool polymerase extension reaction, a semi - random primer, e.g., 5' - (M) k -(X) i -(N) j -3' (where (X) i is a random sequence of nucleotides of length i nucleotides (a set of primers that includes all possible combinations of (X) i ), (N) j is a specific nucleotide (or a series of j nucleotides), and (M) k is a specific nucleotide (or a series of k nucleotides)) is hybridized to the 3' end of the probability barcode tethered to the beads in a given pool. In this case, different deoxyribonucleotide triphosphates (dNTPs) are added to each pool and incorporated by the polymerase into the tethered oligonucleotide.

[0147] The number of combinatorial barcode reagents conjugated onto or synthesized on a solid carrier may include at least 100, 1,000, 10,000, or 1,000,000 or more combinatorial barcode reagents. The number of combinatorial barcode reagents conjugated onto or synthesized on a solid carrier may include at most 100, 1,000, 10,000, or 1,000,000 or more combinatorial barcode reagents. The number of combinatorial barcode reagents conjugated onto or synthesized on a solid carrier such as beads may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 times or more the number of target nucleic acids in the cell. The number of combinatorial barcode reagents conjugated onto or synthesized on a solid carrier such as beads may be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 times or more the number of target nucleic acids in the cell. At least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% of the probability barcodes are bindable to target nucleic acids. At most 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% of the probability barcodes are bindable to target nucleic acids. At least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or more different target nucleic acids can be captured by the combinatorial barcode reagent on the solid carrier. At most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or more different target nucleic acids can be captured by combinatorial barcoding reagents on a solid carrier.

[0148] How to add barcodes to multiple partitions This disclosure provides a method for barcoded multiple partitions using a set of combinatorial barcodes. In some embodiments, the multiple partitions may be multiple microwells in a microwell array. In some embodiments, the number of microwells in a microwell array may be at least 96, at least 384, at least 1,536, at least 5,000, at least 10,000, at least 25,000, at least 50,000, at least 75,000, at least 100,000, at least 500,000, at least 1,000,000, or at least 5,000,000 or more. In some embodiments, each microwell array may include an array barcode that is identifiable with the array barcodes of other microwell arrays.

[0149] In some embodiments, the method may include the step of introducing several component barcodes into each of a plurality of partitions, under the condition that the component barcodes are selected from the set of component barcodes disclosed herein. In some embodiments, the component barcodes in each of the plurality of partitions may be concatenated with one another to form a combinatorial barcode. The component barcodes may be concatenated with or without the presence of the target polynucleotide of the sample.

[0150] As shown in Figure 3, combinatorial barcode reagents 305, 310, 315, and 320 can be dispensed into any specific microwell 325 (e.g., an array of microwells 330). Dispensing can be carried out by methods such as pipetting, pinspotting, or non-contact printing. Combinatorial barcode reagents can also be randomly added to wells using delivery methods such as beads or other particles. For example, an inkjet printhead can dispense combinatorial barcode reagents into each well. Barcoding reagents can be dispensed or formed in individual partitions. Exemplary partitions, but not limited to, include tubes, wells in microtiter plates, flow cells, hollow fibers, microwells on slides, hollow closed or open particles, closed droplets, particles such as beads or other solid, liquid, or gel particles, or other physical partitions used for separating samples. Physical separation can also be achieved by spatially separating individual samples on a two-dimensional surface or substrate within a three-dimensional solid or scaffold.

[0151] Since combinatorial barcode reagents can be designed in this manner, each subunit code sequence in each subunit code section can be combined in a pre-programmed order using a linker. The final code in each microwell 325 can be unique. For any given combination of combinatorial barcode reagents, the position of the microwell can be determined. Thus, combinatorial barcode reagents can be used to determine the spatial position of a sample in a microwell 325 on a microwell array 330.

[0152] This disclosure provides compositions and methods for generating a large repertoire of barcoding reagent diversity (e.g., combinatorial barcoding) by using only a small set of component barcoding reagents. Combinatorial barcoding methods can be combined with probabilistic barcoding methods. In some cases, it is possible to process hundreds of thousands to millions or more samples in parallel. For example, the number of samples that can be processed in parallel may be at least 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, or 900,000 samples. The number of samples that can be processed in parallel may be at most 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, or 900,000 samples. The number of samples that can be processed in parallel may be at least 1,000,000, 2,000,000, 3,000,000, 4,000,000, 5,000,000, 600,000, 700,000, 800,000, or 900,000 samples. The number of samples that can be processed in parallel can be at most 1 million, 2 million, 3 million, 4 million, 5 million, 6 million, 7 million, 8 million, 9 million, or 10 million samples or more.

[0153] A method for simultaneously barcode-tagging a large number (thousands to millions) of individual nucleic acid samples is disclosed herein. A barcode tag may consist of one or more random or predetermined / known nucleic acid sequence strings (barcodes). A barcode may include synthetic oligonucleotides and / or natural or non-natural DNA fragments. For example, a barcode may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides long or more, or at least such a length.

[0154] In addition to various DNA sequences, combinatorial barcodes can be composed of various materials or arise in various forms. For example, a component barcode may be a combination of beads and gel particles, or a combination of a second oligonucleotide on a bead and a first oligonucleotide in solution, or a combination of a first oligonucleotide immobilized on a substrate in a microwell and a second oligonucleotide present in solution in a microwell or on solid particles. Barcodes can also be embedded in hydrogels or similar materials that can act as sponge-like scaffolds for the purpose of localizing biological samples and nucleic acids to specific locations. For example, it is possible to lyse cells on a thin tissue section and capture the released nucleic acid content at a predetermined location. The captured nucleic acid may incorporate a position-specific barcode or combinatorial barcode and, if desired, be further replicated and detected at a predetermined location, or released and isolated for further characterization by methods such as next-generation sequencing.

[0155] The sample can be dispensed into the microwells. In some embodiments, the sample may be a single cell. In some embodiments, the sample may be a tissue section. The sample can be dispensed before dispensing the combinatorial barcode reagent. The sample can be dispensed after dispensing the combinatorial barcode reagent. The sample can be dispensed using a dispensing (e.g., isolation) device. Exemplary dispensing / isolation devices include, but are not limited to, flow cytometers, needle arrays, and microinjectors. The microwell array can be mounted on an isolation / dispensing device so that the sample can be isolated / dispensed directly into the wells of the microwell array.

[0156] Samples containing the barcodes of this disclosure (i.e., samples subjected to the barcoding method of this disclosure) can be used for downstream applications such as sequencing. After DNA sequencing, in addition to the obtained sample nucleic acid sequence, each sequenced read may have a sample barcode string that allows, for example, computer processing to assign that read to a specific sample in a pool of samples.

[0157] Samples can be dispensed into the wells of the substrate using a non-random method. The position of each sample dispensed / isolated in the microwell array may be known. In this way, the distribution does not have to follow randomness and / or probability. Samples can be dispensed into the wells of the substrate using a non-Poisson method. The Poisson distribution, or curve, is a discrete probability distribution that represents the probability of several events occurring in a given period of time, given that the events occur at a known mean rate and are independent of each other. The Poisson distribution formula is as follows: f(k;λ)=(e -λ λ k The formula is ( / k!), where k is the number of event occurrences and λ is a positive real number representing the expected number of occurrences at a given interval. A non-Poisson distribution can be any distribution that does not follow the Poisson formula.

[0158] Microwells containing nucleic acids, single cells, single-cell nucleic acids, combinatorial barcode reagents, and reagents necessary for primer extension and amplification (e.g., polymerase dNTP, buffer), or any combination thereof, can be sealed. Sealing of microwells can be useful in preventing cross-hybridization of target nucleic acids between adjacent microwells. Microwells can be sealed with tape and / or any adhesive. The sealant can be transparent. The sealant can be opaque. The sealant can be light-transmitting, allowing for visual detection of the contents of the sealed microwells (e.g., UV-Vis, fluorescence). Nucleic acids in partitions (e.g., microwells of an array) (e.g., those of a sample) can be associated (e.g., hybridized) with combinatorial barcode reagents dispensed within the microwells.

[0159] The combinatorial barcode reagent and target polynucleotide in the wells can be extended and / or amplified to produce transcripts and / or amplicons containing the sequence of the combinatorial barcode reagent. Amplification can be performed by methods including, but is not limited to, PCR, primer extension, reverse transcription, isothermal amplification, linear amplification, multiple substitution amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand substitution amplification (SDA), real-time SDA, rolling circle amplification, or circle-circle amplification, RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, and assembly PCR, endpoint PCR, and suppression PCR, or any combination thereof. In some cases, amplification and / or extension is performed by isothermal amplification. In some cases, amplification and / or extension is performed by multiple strand substitution.

[0160] Microwell contents can be pooled after extension and / or amplification. Pooling the contents can reduce errors during sample preparation for downstream processes. Microwell samples can be pooled in one container (e.g., a tube). Microwell samples can be pooled in two or more containers (e.g., tubes for sample / plate indexing).

[0161] Pooled microwell samples are molecularly manipulable. For example, a pooled sample can be subjected to one or more rounds of amplification. Pooled samples can be subjected to purification steps (e.g., using Ampure beads, size selection, gel / column filtration, removal of rRNA or any other RNA impurities, enzymatic degradation of impurities, pulsed gel electrophoresis, and precipitation with a sucrose gradient or cesium chloride gradient, and size exclusion chromatography (gel permeation chromatography)).

[0162] Pooled microwell samples can be prepared for sequencing. Sequencing preparations may involve sequencing (e.g., flow cell adapter) by either adapter ligation (e.g., TA ligation) or primer extension (e.g., in this case the primers contain the sequence of the flow cell adapter). Methods for ligating adapters to nucleic acid fragments are well known. Adapters can be double-stranded, single-stranded, or partially single-stranded. In some embodiments, the adapter is formed from two oligonucleotides having complementary regions, e.g., about 10–30 or about 15–40 bases of complete complementarity, such that the two oligonucleotides form a double-stranded region when hybridized integrally. Optionally, one or both oligonucleotides may have regions that are not complementary to the other oligonucleotide, and may form single-stranded overhangs at one or both ends of the adapter. The single-stranded overhangs may be about 1–8 or about 2–4 bases. The overhangs are complementary to those formed by restriction enzyme cleavage and can promote "adherent end" ligation. The adapter may include other features such as primer binding sites and restriction sites. In some embodiments, the restriction site is for an IIS-type restriction enzyme or another enzyme that cleaves outside the recognition sequence, such as EcoP151. The pooled sample is sequenceable (e.g., by the method described herein).

[0163] Sequencing of combinatorial barcoded nucleic acids may include the step of sequencing at least a portion of the combinatorial barcode (for example, the sequence of the combinatorial barcode reagent used to construct the combinatorial barcode). Sequencing may include the step of sequencing at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% of the combinatorial barcode of the target polynucleotide. Sequencing may include the step of sequencing as many as 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% of the combinatorial barcode of the target polynucleotide. Sequencing may include the step of sequencing at least a portion of the target polynucleotide. Sequencing may include the step of sequencing at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% of the target polynucleotide. Sequencing may include the step of sequencing at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% of the target polynucleotide. Sequencing may include the step of sequencing at least a portion of the combinatorial barcode and at least a portion of the target polynucleotide. Sequencing may include the step of sequencing the conjugate between the combinatorial barcode and the target polynucleotide. Sequencing may include paired-end sequencing. Paired-end sequencing may allow determination of two or more reads of sequences at two locations on a single polynucleotide double strand.

[0164] After DNA sequencing, in addition to the obtained sample nucleic acid sequence, each sequenced read may have a sample barcode string and / or target nucleic acid sequence that allows the read to be assigned to a specific sample in a pool of samples, for example, using computer processing. For example, reads containing combinatorial barcode sequences in a first microwell can be binned together, while reads containing combinatorial barcode sequences in a second microwell can be binned together. Reads for each microwell (e.g., bin) can be analyzed and / or processed. For example, if the starting material was fragmented chromosomal DNA, the reads can be processed to form contigs and / or haplotype phasing can be performed as described in PCT International Application No. PCT / US2016 / 22712 (the contents of which are incorporated in whole by reference).

[0165] For example, contigs (contiguous sequences) can be formed from the short-range characteristic data of the reads. Reads can be binned. In every bin, the short-range probe signatures of the reads in that bin can be retrieved and recorded. By using conventional contigging analysis on a subset of the binned reads, the order and distance of the contigated reads in a bin can be determined using the short-range signatures. Larger contigs can be formed by the orientation and connection of contigs in adjacent bins.

[0166] In some embodiments, reads are binned according to short-range probing. After an initial comparison of short-range probe signatures, clusters of reads with high confidence are formed. The order and distance of reads within the clusters are then determined. Long-range probe information for all read clusters is retrieved and recorded, and a composite long-range score is generated for the clusters. This composite can be generated for each entry by taking the maximum (or other arithmetic combination) of the (read) comparison scores across all reads in the cluster. The result is a binning of clusters relative to the genome, and this global positioning information can be used to determine neighboring clusters. The contigs of neighboring binned clusters can then be oriented and connected to form larger contigs.

[0167] Barcoding of multiple DNA targets Several embodiments disclosed herein provide a method for barcoding multiple DNA targets in multiple partitions. Multiple partitions can be barcoded using the combinatorial barcodes disclosed herein. For example, each partition can be barcoded with a unique combination of component barcodes selected from a set of component barcodes. In some embodiments, each combinatorial barcode in the multiple partitions includes a target-specific region that binds to one or more DNA targets in the partition. Multiple DNA targets, such as genomic DNA fragments, can be introduced into the multiple partitions. In some embodiments, each partition contains one or fewer DNA targets. In some embodiments, each partition contains one or fewer DNA targets from the same chromosome. In some embodiments, each partition may contain, on average, 1 or fewer, 2 or fewer, 3 or fewer, 5 or fewer, 10 or fewer, or 100 or fewer DNA targets. In some embodiments, DNA targets are introduced into the partitions after the component barcodes have been introduced into the partitions. In some embodiments, DNA targets are introduced into the partitions before the component barcodes have been introduced into the partitions. In some embodiments, one or more component barcodes in each partition hybridize to one or more DNA targets in the partitions. In some embodiments, component barcodes hybridized to a DNA target can be used as primers for an extension reaction to generate an extension product containing the DNA target and a combinatorial barcode. In some embodiments, the extension product can be amplified and sequenced.

[0168] Figure 4 shows an exemplary embodiment of the combinatorial barcoding method of the present disclosure for generating contigs. Nucleic acids (e.g., chromosomes) are fragmentable. Fragmentation can be performed by any method, such as sonication, shearing, enzymatic fragmentation, etc. The fragments can be 10, 50, 100, 150, 200, 250, 300, 350, or 400 kilobases long or more or at least such lengths. The fragments can be at most 10, 50, 100, 150, 200, 250, 300, 350, or 400 kilobases long or more. The fragmentation step can generate at least 1×10 3 、1×10 4 、1×10 5 、1×10 6 、1×10 7 、or 1×10 8 fragments or more. The fragmentation step can generate at most 1×10 3 、1×10 4 、1×10 5 、1×10 6 、1×10 7 、or 1×10 8Fragments or more can be generated. The fragments isolated in the wells may be approximately 1-20, 1-15, 1-10, 1-5, 5-20, 5-15, 5-10, 10-15, or 10-20 fragments / well. The fragments isolated in the wells may overlap. The fragments isolated in the wells may not overlap. Nucleic acid fragment 405 can be isolated in well 410 of the microwell array 415. Some or all of the wells of the microwell array may contain one or more combinatorial barcode reagents 420 and 425 (e.g., a known grouping of combinatorial barcode reagents generates a unique combinatorial barcode). The microwells are sealable. The contents of the microwells can be extended / amplified to generate a transcript 430 containing a combinatorial barcode (e.g., by primer extension and / or multi-strand substitution). Transcripts / amplicons 430 containing combinatorial barcodes are sequenceable. Sequencing analysis may involve binning reads based on combinatorial barcode sequences corresponding to the same microwells. Reads can be generated and contigs can be used for haplotype phasing (e.g., parent SNPs).

[0169] Single-cell sequencing Several embodiments disclosed herein provide a method for single-cell sequencing. Multiple partitions can be barcoded using combinatorial barcodes disclosed herein. For example, each partition can be barcoded with a unique combination of component barcodes selected from a set of component barcodes. In some embodiments, each combinatorial barcode for each of the multiple partitions includes a target-specific region that binds to one or more target polynucleotides of a single cell in the partition. In some embodiments, the target polynucleotides include DNA. In some embodiments, the target polynucleotides include RNA.

[0170] Multiple single cells can be introduced into multiple partitions, such as microwells in a microwell array. In some embodiments, the cells can be enriched with respect to desired properties using a flow cytometer. In some embodiments, the step of introducing multiple single cells into the microwells of a microwell array may include the step of depositing multiple cells into the microwells of the microwell array by flow cytometry. The step of depositing multiple single cells into the microwells of a microwell array by flow cytometry may include the step of depositing single cells into the microwells of the microwell array at one time using a flow cytometer. In some embodiments, the multiple partitions consist of more than 1536 microwells. In some embodiments, the multiple partitions are in the form of a microwell array of 5 micrometers or larger. In some embodiments, the multiple partitions are dimensionally defined by spatial positioning or the like.

[0171] Figure 5 shows an exemplary method for single-cell sequencing. The barcoded sample may contain a single cell and / or single-cell nucleic acid. A single cell 505 can be isolated in well 510 of a microwell array 515. In some embodiments, at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% of the wells may contain single cells. In some embodiments, at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% of the wells may contain single cells. Cells can be distributed into the wells non-Poissonianly. Cells can be distributed into the wells non-randomly. Cells can be distributed into the wells Poissonianly. Cells can be distributed into the wells randomly.

[0172] Cells can release their nucleic acid content by lysis. Wells can be sealed before lysis. The nucleic acid content of cells can be associated with one or more combinatorial barcode reagents 520 / 525 in the microwells (e.g., generating a unique combinatorial barcode by known grouping of combinatorial barcode reagents). The contents of the microwells can generate transcripts 530 containing combinatorial barcodes by extension / amplification (e.g., by primer extension and / or multistrand substitution). Transcripts / amplicons 530 containing combinatorial barcodes are sequenceable. Sequencing analysis may include the step of binning reads based on combinatorial barcode sequences corresponding to the same microwells. Reads are analyzable for single-cell gene expression analysis. Reads are usable for analyzing cells (e.g., to determine cell type, to diagnose cells as part of a disorder or disease, to genotype cells, to determine or predict cell response to a therapy or regimen, and / or to measure changes in the cell expression profile over time).

[0173] In some cases, the methods of the present disclosure may include probabilistic barcoding and combinatorial barcoding. For example, as shown in Figure 6, cells 605 can be non-Poissonianly distributed into wells 610 in a microwell array 615 such that each well contains a single cell. The cells are lysable. The nucleic acids 620 of the cells can be contacted by a plurality of probabilistic barcodes 625, which include oligo dT of the present disclosure that can contact a target (e.g., mRNA), as well as cell labels and molecular labels. The probabilistic barcodes can generate probabilistic barcoded cDNA 630 by reverse transcription. The probabilistic barcoded cDNA can be contacted by one or more combinatorial barcoding reagents 635 / 640 that can perform a second strand synthesis. One or more combinatorial barcoding reagents 635 / 640 for the second strand synthesis may include a universal sequence (e.g., a sequencing primer binding site, i.e., Illumina read 1). The second chain synthesis may yield nucleic acid molecule 645 with both a probabilistic barcode and a combinatorial barcode. This nucleic acid molecule 645, with both a probabilistic barcode and a combinatorial barcode, can be sequenced by downstream library preparation methods (e.g., pooling, sequencing adapter addition). The probabilistic barcode can be used to count the number of target nucleic acids 620 in sample 605. The combinatorial barcode can be used to identify well 610 of the microwell array 615 that contained the cells.

[0174] Binatorial barcoding may be included. For example, as shown in Figure 7, cells 705 can be non-Poissonianly distributed into wells 710 in a microwell array 715 such that each well contains a single cell. The cells are lysable. The nucleic acids 720 of the cells can be contacted with one or more 3' combinatorial barcoding reagents 725 that can contact a target (e.g., mRNA). One or more 3' combinatorial barcoding reagents 725 generate 3' combinatorial barcoded cDNA 730 by reverse transcription. The 3' combinatorial barcoded cDNA 730 can be contacted with one or more 5' combinatorial barcoding reagents 740 that can perform a second strand synthesis. One or more 5' combinatorial barcoding reagents 740 for the second strand synthesis may contain a universal sequence (e.g., a sequencing primer binding site, i.e., Illumina read 1). The second chain synthesis may yield a nucleic acid molecule 745 with a viper-tite combinatorial barcode. This viper-tite combinatorial barcode nucleic acid molecule 745 can be sequenced by downstream library preparation methods (e.g., pooling, sequencing adapter addition). The viper-tite combinatorial barcode can identify well 710 of the microwell array 715 that contained the cells.

[0175] The methods disclosed herein enable the comparison of cells and / or nucleic acid transcripts at various points in time. The methods disclosed herein can be used to diagnose subjects or to compare cells / nucleic acids of diseased subjects with those of healthy subjects. For example, each well of a microwell array may contain cells from different subjects.

[0176] The methods disclosed herein can be used, for example, in pharmaceutical or clinical trial testing. For example, microwells of a substrate containing a combinatorial barcode reagent may contain a modulator reagent. The modulator reagent may contain a drug, therapeutic agent, protein, RNA, liposome, lipid, or any therapeutic molecule (e.g., a clinical trial molecule, an FDA-approved molecule, or an FDA-unapproved molecule). The modulator agent can be introduced into the well before the isolation of the sample into the well. The modulator agent can be introduced into the well after the isolation of the sample into the well (e.g., the modulator agent can come into contact with the sample in the well). As an example, each well in a microwell array may contain cells from various subjects. The cells can be treated with the modulator agent. The sequencing results can be used to determine how the modulator agent affected gene expression in a sample (e.g., a single cell). The sequencing results can be used to determine which subjects in a subject population may be affected by the modulator agent and / or to diagnose subjects in a population.

[0177] Probability barcoding method This disclosure provides methods for combinatorial barcoding and / or probabilistic barcoding of samples. The combinatorial barcoding method and the probabilistic barcoding method are combinatorial. The method may include the steps of positioning a probabilistic barcode in close proximity to the sample, lysing the sample, associating identifiable targets with the probabilistic barcode, amplifying the targets, and / or digitally counting the targets. The method may further include the steps of analyzing and / or visualizing information obtained from spatial labels on the probabilistic barcode. A sample (e.g., a sample section, thin slice, or cell) can come into contact with a solid carrier containing a probabilistic barcode. Targets in the sample can be associated with the probabilistic barcode. The solid carrier is collectible. cDNA synthesis can be performed on the solid carrier. cDNA synthesis can be performed away from the solid carrier. cDNA synthesis can generate a target-barcode molecule by incorporating labeling information from the labels in the probabilistic barcode into the newly synthesized cDNA target molecule. The target-barcode molecule can be amplified using PCR. The sequence of target and probability barcode labels on a target-barcode molecule can be determined by sequencing.

[0178] Contact between sample and probability barcode This disclosure provides a method for distributing a sample (e.g., cells) into multiple wells of a substrate under conditions where the wells contain combinatorial barcode reagents and the distribution is performed non-Poissonianly. The distribution / isolation of cells into the wells can be non-random (e.g., cells are specifically sorted to particular positions in the array). In this way, the distribution of cells can be considered non-Poissonian.

[0179] For example, a sample containing cells, organs, or tissue flakes can be in contact with a probability barcode. The solid carrier may be floating. The solid carrier can be embedded in a semi-solid or solid array. The probability barcode can be associated with the solid carrier. The probability barcode may be individual nucleotides. The probability barcode can be associated with a substrate. If the probability barcode is in very close proximity to the target, the target can hybridize to the probability barcode. The probability barcode can be in contact with each identifiable target in a non-depletion ratio so that each identifiable target can be associated with an identifiable probability barcode of this disclosure. To ensure efficient association between the target and the probability barcode, the target can be crosslinked to the probability barcode.

[0180] The probability that two identifiable targets in a sample can contact the same unique probability barcode is 10. -6 , 10 -5 , 10 -4 , 10 -3 , 10 -2 , or 10 -1 Or it could be more or at least that number. The probability that two identifiable targets in the sample can access the same unique probability barcode is at most 10 -6 , 10 -5 , 10 -4 , 10 -3 , 10 -2 , or 10 -1 Or it could be more. The probability that two targets of the same gene in the same cell can access the same probability barcode is 10 -6 , 10 -5 , 10 -4 , 10 -3 , 10 -2 , or 10 -1 Or it could be more or at least that number. The probability that two targets of the same gene in the same cell can access the same probability barcode is at most 10 -6 , 10 -5 , 10 -4 , 10 -3 , 10 -2 , or 10 -1 Or it could be even more.

[0181] In some cases, cells of a cell population can be separated (e.g., isolated) within wells of the substrate of this disclosure. The cell population can be diluted before separation. The cell population can be diluted such that at least 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100% of the wells of the substrate contain single cells. The cell population can be diluted such that at least 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100% of the wells of the substrate contain single cells. The cell population can be diluted so that the number of cells in the diluted population is 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100% or at least such number of cells in the diluted population, relative to the number of wells on the substrate. In some cases, the cell population is diluted so that the number of cells is approximately 10% of the number of wells on the substrate.

[0182] The distribution of single cells into the wells of the substrate may follow a Poisson distribution. For example, the probability that a well of the substrate contains two or more cells may be at least 0.1, 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10% or more. The distribution of single cells into the wells of the substrate may be random. The distribution of single cells into the wells of the substrate may be non-random. Cells can be separated so that each well of the substrate contains only one cell.

[0183] Cell lysis After the distribution of cells and probabilistic barcodes, the cells can be lysed to release the target molecules. Cell lysis can be achieved by any of the following means, for example, by chemical or biochemical means, by osmotic shock, or by thermal lysis, mechanical lysis, or optical lysis. Cells can be lysed by adding a cell lysis buffer containing a surfactant (e.g., SDS, Li dodecyl sulfate, Triton X-100, Tween-20, or NP-40), an organic solvent (e.g., methanol or acetone), or a digestive enzyme (e.g., proteinase K, pepsin, or trypsin), or any combination thereof. To improve the association between the target and the probabilistic barcode, the diffusion rate of the target molecules can be altered, for example, by decreasing the temperature and / or increasing the viscosity of the lysate.

[0184] Attaching probabilistic barcodes to target nucleic acid molecules Following cell lysis and the release of nucleic acid molecules therefrom, the nucleic acid molecules can be randomly associated with probabilistic barcodes on a colocalized solid carrier. This association may involve, for example, hybridization of the target recognition region of the probabilistic barcode to a complementary portion of the target nucleic acid molecule (e.g., the oligo-dT of the probabilistic barcode can interact with the poly-A tail of the target). The assay conditions used for hybridization (e.g., buffer pH, ionic strength, temperature, etc.) can be selected to facilitate the formation of specific stable hybrids.

[0185] The binding may further involve ligating a target recognition region of the probabilistic barcode with a portion of the target nucleic acid molecule. For example, the target binding region may include a nucleic acid sequence that can be specifically hybridized to a restriction site overhang (e.g., an EcoRI attachment end overhang). The assay procedure may further include treating the target nucleic acid with a restriction enzyme (e.g., EcoRI) to generate a restriction site overhang. The probabilistic barcode can then be ligated to any nucleic acid molecule containing a sequence complementary to the restriction site overhang. A ligase (e.g., T4 DNA ligase) may be used to ligate the two fragments.

[0186] Labeled targets (e.g., target-barcode molecules) from multiple cells (or multiple samples) can be subsequently pooled, for example, by recovering beads to which the probabilistic barcodes and / or target-barcode molecules are bound. Recovery of a solid-support-based collection of bound target-barcode molecules can be achieved using magnetic beads and an externally applied magnetic field. After pooling the target-barcode molecules, all further processing can be carried out in a single reaction vessel. Further processing may include, for example, reverse transcription, amplification, cleavage, dissociation, and / or nucleic acid extension reactions. Further processing reactions can be carried out in microwells, i.e., without first pooling the labeled target nucleic acid molecules from multiple cells.

[0187] Reverse transcription This disclosure provides a method for generating a stochastic target-barcode conjugate using reverse transcription. The stochastic target-barcode conjugate may include a stochastic barcode and a complementary sequence of all or part of the target nucleic acid (i.e., a cDNA molecule with a stochastic barcode). Reverse transcription of the associated RNA molecule can be performed by adding a reverse transcription primer along with reverse transcriptase. The reverse transcription primer may be an oligo-dT primer, a random hexanucleotide primer, or a target-specific oligonucleotide primer. Oligo-dT primers may be 12–18 nucleotides long and can bind to the endogenous poly(A) tail at the 3' end of mammalian mRNA. Random hexanucleotide primers can bind to mRNA at various complementary sites. Target-specific oligonucleotide primers typically selectively prime the target mRNA.

[0188] amplification Nucleic acid amplification reactions may be performed one or more times to generate multiple copies of a labeled target nucleic acid molecule. Amplification may be performed in a multiplexing manner under conditions in which multiple target nucleic acid sequences are amplified simultaneously. Amplification reactions may be used to add sequencing adapters to nucleic acid molecules. Amplification reactions may include a step of amplifying at least a portion of sample labels, if present. Amplification reactions may include a step of amplifying at least a portion of cell and / or molecular labels. Amplification reactions may include a step of amplifying at least a portion of sample tags, cell labels, spatial labels, molecular labels, target nucleic acids, or combinations thereof. The amplification reaction may include a step of amplifying at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 100% of a plurality of nucleic acids. The method may further include a step of performing one or more cDNA synthesis reactions to produce one or more cDNA copies of a target-barcode molecule, including sample labeling, cell labeling, spatial labeling, and / or molecular labeling.

[0189] In some embodiments, amplification may be performed using polymerase chain reaction (PCR). As used herein, PCR may mean a reaction that amplifies a specific DNA sequence in vitro by simultaneous primer extension of the complementary strand of DNA. As used herein, PCR may encompass derivatives of the reaction, such as, but not limited to, RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, and assembly PCR.

[0190] Amplification of labeled nucleic acids may include non-PCR-based methods. Examples of non-PCR-based methods, but are not limited to, multiple substitution amplification (MDA), transcription-mediated amplification (TMA), whole transcriptome amplification (WTA), whole genome amplification (WGA), nucleic acid sequence-based amplification (NASBA), strand substitution amplification (SDA), real-time SDA, rolling circle amplification, or circle-circle amplification. Other non-PCR-based amplification methods include DNA-dependent RNA polymerase-driven RNA transcription amplification or multiple cycles of RNA-directed DNA synthesis and transcription for amplifying DNA or RNA targets, ligase chain reaction (LCR), and Qβ replicase (Qβ) methods, the use of palindromic probes, strand substitution amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods in which primers are hybridized to nucleic acid sequences and the resulting double strands are cleaved before extension and amplification, strand substitution amplification using nucleic acid polymerases lacking 5' exonuclease activity, rolling circle amplification, and branched extension amplification (RAM). In some cases, amplification can produce cyclized transcripts.

[0191] Suppression PCR can be used in the amplification method of this disclosure. Suppression PCR may mean that molecules below a certain size are selectively excluded due to inefficient amplification resulting from the terminal reverse repeat flanking when the primers used for amplification correspond to all or part of the repeat. This may be due to the equilibrium between productive PCR primer annealing and unproductive self-annealing of the complementary ends of the fragment. For a given size of flanking terminal reverse repeat, the shorter the insert, the stronger the suppression effect, and vice versa. Similarly, for a given insert size, the longer the terminal reverse repeat, the stronger the suppression effect.

[0192] Suppression PCR can utilize adapters ligated to the ends of DNA fragments before PCR amplification. After thawing and annealing, single-stranded DNA fragments with self-complementary adapters at the 5' and 3' ends of the strand can generate suppressive "tennis racket" shaped structures that inhibit the amplification of the fragment during PCR.

[0193] In some cases, the methods disclosed herein further include the step of carrying out a polymerase chain reaction on a labeled nucleic acid (e.g., labeled RNA, labeled DNA, labeled cDNA) to produce a probabilistically labeled amplicon. The labeled amplicon may be a double-stranded molecule. The double-stranded molecule may include a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized to a DNA molecule. One or both strands of the double-stranded molecule may include sample labeling, spatial labeling, cellular labeling, and / or molecular labeling. The probabilistically labeled amplicon may be a single-stranded molecule. The single-stranded molecule may include DNA, RNA, or a combination thereof. The nucleic acids of this disclosure may include synthetic nucleic acids or modified nucleic acids.

[0194] Amplification may involve the use of one or more non-natural nucleotides. Non-natural nucleotides may include photounstable or triggering nucleotides. Examples of non-natural nucleotides, but are not limited to, peptide nucleic acids (PNA), morpholino nucleic acids, and locked nucleic acids (LNA), as well as glycol nucleic acids (GNA) and threose nucleic acids (TNA). Non-natural nucleotides may be added to one or more cycles of the amplification reaction. The addition of non-natural nucleotides may be used to identify the product at a particular cycle or point in time of the amplification reaction.

[0195] A step involving one or more amplification reactions may include the use of one or more primers. One or more primers may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides or more. One or more primers may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides or more. One or more primers may contain less than 12 to 15 nucleotides. One or more primers may anneal to at least a portion of multiple stochastic-labeled targets. One or more primers may anneal to the 3' or 5' ends of multiple stochastic-labeled targets. One or more primers may anneal to the internal regions of multiple stochastic-labeled targets. The internal region may consist of at least approximately 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of multiple probabilistically labeled targets. One or more primers may contain a constant panel of primers. One or more primers may include at least one custom primer. One or more primers may include at least one control primer. One or more primers may include at least one gene-specific primer.

[0196] One or more primers may include any universal primer of this disclosure. Universal primers may anneal to a universal primer binding site. One or more custom primers may anneal to a first sample label, a second sample label, a spatial label, a cellular label, a molecular label, a target, or any combination thereof. One or more primers may include universal primers and custom primers. Custom primers may be designed to amplify one or more targets. Targets may include a subset of all nucleic acids in one or more samples. Targets may include a subset of all probabilistic labeled targets in one or more samples. One or more primers may include at least 96 custom primers or more. One or more primers may include at least 960 custom primers or more. One or more primers may include at least 9600 custom primers or more. One or more custom primers may anneal to two or more different labeled nucleic acids. Two or more different labeled nucleic acids may correspond to one or more genes.

[0197] Any amplification scheme can be used in the method of this disclosure. For example, in one scheme, a first round of PCR can amplify a molecule (e.g., conjugated to beads) using a gene-specific primer and a primer for universal Illumina sequencing primer 1 sequence. A second round of PCR can amplify the first PCR product using a nested gene-specific primer flanked by Illumina sequencing primer 2 sequence and a primer for universal Illumina sequencing primer 1 sequence. A third round of PCR adds P5, P7 and a sample index to the PCR product to form an Illumina sequencing library. Sequencing using 150 bp × 2 sequencing can reveal cell labels and molecular indices on read 1, genes on read 2, and a sample index on index 1 read.

[0198] Amplification can be performed in one or more rounds. In some cases, multiple rounds of amplification are performed. Amplification may involve two or more rounds. The first amplification may be an extension off X' that generates a gene-specific region. The second amplification may occur when the sample nucleic acid hybridizes to the newly formed strand.

[0199] In some embodiments, hybridization does not need to be performed at the ends of the nucleic acid molecule. In some embodiments, the target nucleic acid is hybridized and amplified within the intact strand of a longer nucleic acid. For example, a target within a longer section of genomic DNA or mRNA. The target can be more than 50 nt, more than 100 nt, or more than 1000 nt from the ends of a polynucleotide.

[0200] Sequencing The step of determining the number of different probabilistically labeled nucleic acids may include the step of sequencing labeled targets, spatially labeled, molecularly labeled, sample labeled, and cellular labeled products, or any product (e.g., labeled amplicons, labeled cDNA molecules). Amplification targets may be subjected to sequencing. The step of sequencing probabilistically labeled nucleic acids or any product thereof may include the step of performing sequencing reactions to sequence at least some of the sample labeled, spatially labeled, cellular labeled, and molecular labeled products, and / or at least some of the probabilistically labeled targets, their complements, their reverse complements, or any combination thereof.

[0201] The sequencing of nucleic acids (e.g., amplified nucleic acids, labeled nucleic acids, cDNA copies of labeled nucleic acids, etc.) is not limited to synthetic sequencing (SBS), hybridization sequencing (SBH), ligation sequencing (SBL), quantitative progressive fluorescence nucleotide addition sequencing (QIFNAS), stepwise ligation and cleavage, fluorescence resonance energy transfer (FRET), molecular beacons, TaqMan reporter probe digestion, pyrosequencing, and fluorescence in This can be performed using a variety of sequencing methods, including situ sequencing (FISSEQ), FISSEQ beads, wobble sequencing, multiple sequencing, polymerized colony (POLONY) sequencing, nanogrid rolling circle sequencing (ROLONY), and allele-specific oligoligation assays (e.g., oligoligation assay (OLA), single-template molecule OLA using a ligate linear probe and rolling circle amplification (RCA) readout, ligate padlock probe, or single-template molecule OLA using a ligate cyclic padlock probe and rolling circle amplification (RCA) readout).

[0202] In some cases, the process of determining the sequence of a labeled nucleic acid or any of its products includes paired-end sequencing, nanopore sequencing, high-throughput sequencing, shotgun sequencing, dye-terminator sequencing, multiprimer DNA sequencing, primer walking, Sangerdideoxy sequencing, Maxam-Gilbert sequencing, pyrosequencing, true single-molecule sequencing, or any combination thereof. Alternatively, the sequence of a labeled nucleic acid or any of its products may be determined by electron microscopy or a chemisensitized field-effect transistor (chemFET) array.

[0203] High-throughput sequencing methods, such as cyclic array sequencing using platforms like Roche454, Illumina Solexa, ABI-SOLiD, ION Torrent, Complete Genomics, Pacific Bioscience, Helicos, and Polonator, can also be utilized. Sequencing may include MiSeq sequencing. Sequencing may include HiSeq sequencing.

[0204] Probabilistically labeled targets may contain nucleic acids corresponding to approximately 0.01% to approximately 100% of the genes in an organism's genome. For example, by capturing a gene containing a complementary sequence of a sample, it is possible to sequence approximately 0.01% to approximately 100% of the genes in an organism's genome using a target complementary region containing multiple multimers. In some embodiments, the labeled nucleic acid contains nucleic acids corresponding to approximately 0.01% to approximately 100% of the transcripts in an organism's transcriptome. For example, by capturing mRNA from a sample, it is possible to sequence approximately 0.501% to approximately 100% of the transcripts in an organism's transcriptome using a target complementary region containing a poly-T tail.

[0205] The sequencing may include sequencing of at least approximately 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides or base pairs or more of the labeled nucleic acid and / or probability barcode. The sequencing may include sequencing of many approximately 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides or base pairs or more of the labeled nucleic acid and / or probability barcode. The sequencing may include sequencing of at least approximately 200, 300, 400, 500, 600, 700, 800, 900, 1,000 nucleotides or base pairs or more of the labeled nucleic acid and / or probability barcode. The sequencing may include sequencing of approximately 200, 300, 400, 500, 600, 700, 800, 900, 1,000 nucleotides or base pairs or more of labeled nucleic acids and / or probabilistic barcodes. The sequencing may include sequencing of at least approximately 1,500, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, or 10,000 nucleotides or base pairs or more of labeled nucleic acids and / or probabilistic barcodes. Sequencing may involve sequencing of approximately 1,500, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, or 10,000 nucleotides or base pairs or more of labeled nucleic acids and / or probabilistic barcodes.

[0206] The sequencing may include at least approximately 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 sequencing reads / runs or more. The sequencing may include at most approximately 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 sequencing reads / runs or more. In some cases, the sequencing may include at least approximately 1,500, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, or 10,000 sequencing reads / runs or more. In some cases, sequencing may include at most approximately 1,500, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, or 10,000 sequencing reads / runs or more. Sequencing may include at least 0.1 billion, 0.5 billion, 100 million, 150 million, 200 million, 250 million, 300 million, 350 million, 400 million, 450 million, 500 million, 550 million, 600 million, 650 million, 700 million, 750 million, 800 million, 850 million, 900 million, 950 million, or 1 billion sequencing reads / runs or more. Sequencing may include at most 0.1 billion, 0.5 billion, 100 million, 150 million, 200 million, 250 million, 300 million, 350 million, 400 million, 450 million, 500 million, 550 million, 600 million, 650 million, 700 million, 750 million, 800 million, 850 million, 900 million, 950 million, or 1 billion sequencing reads / runs or more. Sequencing as a whole may include at least 100 million, 200 million, 300 million, 400 million, 500 million, 600 million, 700 million, 800 million, 900 million, 1 billion, 1.1 billion, 1.2 billion, 1.3 billion, 1.4 billion, 1.5 billion, 1.6 billion, 2 billion, 3 billion, 4 billion, or 5 billion sequencing reads / runs or more. Sequencing may, in whole, include at most 100 million, 200 million, 300 million, 400 million, 500 million, 600 million, 700 million, 800 million, 900 million, 1 billion, 1.1 billion, 1.2 billion, 1.3 billion, 1.4 billion, 1.5 billion, 1.6 billion, 2 billion, 3 billion, 4 billion, or 5 billion sequencing reads / runs or more. Sequencing may include approximately 1,600,000,000 sequencing reads / runs or less. Sequencing may include approximately 200,000,000 reads / runs or less.

[0207] Spatial barcoding methods Several embodiments disclosed herein provide methods for spatially barcoding target nucleic acids of a sample. In some embodiments, multiple oligonucleotides are immobilized on a substrate such as a slide. In some embodiments, the substrate may be coated with a polymer, matrix, hydrogel, needle array device, antibody, or any combination thereof. In some embodiments, the oligonucleotides include target-specific regions such as oligo(dT) sequences, gene-specific sequences, or random multimers.

[0208] For example, a sample including slices, cell monolayers, fixed cells, and tissue sections can be in contact with oligonucleotides on a substrate. Cells may include one or more cell types. For example, cells may be brain cells, cardiac cells, cancer cells, circulating tumor cells, organocytes, epithelial cells, metastatic cells, benign cells, primary cells, circulating cells, or any combination thereof. Cells can be in contact by gravity flow under conditions where the cells can remain still and form a monolayer. A sample may be a tissue thin section. The thin section can be placed on a substrate. A sample may be one-dimensional (e.g., forming a planar surface). A sample (e.g., cells) can be spread across the entire substrate by, for example, growing / culturing cells on the substrate.

[0209] Cell lysis After the distribution of cells and probabilistic barcodes, the cells can be lysed to release the target molecules. Cell lysis can be achieved by any of the following means, for example, by chemical or biochemical means, by osmotic shock, or by thermal lysis, mechanical lysis, or optical lysis. Cells can be lysed by adding a cell lysis buffer containing a surfactant (e.g., SDS, Li dodecyl sulfate, Triton X-100, Tween-20, or NP-40), an organic solvent (e.g., methanol or acetone), or a digestive enzyme (e.g., proteinase K, pepsin, or trypsin), or any combination thereof. To improve the association between the target and the probabilistic barcode, the diffusion rate of the target molecules can be altered, for example, by decreasing the temperature and / or increasing the viscosity of the lysate.

[0210] In some embodiments, the sample can be dissolved using filter paper. The filter paper can be immersed in a lysis buffer. The filter paper can be applied to the sample under pressure that can promote the dissolution of the sample and the hybridization of the sample's target into the substrate.

[0211] In some embodiments, dissolution can be carried out by mechanical dissolution, thermal dissolution, optical dissolution, and / or chemical dissolution. Chemical dissolution may involve the use of digestive enzymes such as proteinase K, pepsin, and trypsin. Dissolution can be carried out by adding a dissolution buffer to the substrate. The dissolution buffer may contain Tris-HCl. The dissolution buffer may contain at least about 0.01, 0.05, 0.1, 0.5, or 1 M or more of Tris-HCl. The dissolution buffer may contain at most about 0.01, 0.05, 0.1, 0.5, or 1 M or more of Tris-HCl. The dissolution buffer may contain about 0.1 M Tris-HCl. The pH of the dissolution buffer may be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more. The pH of the dissolution buffer may be at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more. In some embodiments, the pH of the dissolution buffer is about 7.5. The lysis buffer may contain a salt (e.g., LiCl). The concentration of the salt in the lysis buffer may be at least about 0.1, 0.5, or 1 M or higher. The concentration of the salt in the lysis buffer may be at most about 0.1, 0.5, or 1 M or higher. In some embodiments, the concentration of the salt in the lysis buffer is about 0.5 M. The lysis buffer may contain a surfactant (e.g., SDS, Li dodecyl sulfate, Triton X, Tween, NP-40). The concentration of the surfactant in the lysis buffer may be at least about 0.0001, 0.0005, 0.001, 0.005, 0.01, 0.05, 0.1, 0.5, 1, 2, 3, 4, 5, 6, or 7% or higher. The concentration of the surfactant in the lysis buffer may be at most about 0.0001, 0.0005, 0.001, 0.005, 0.01, 0.05, 0.1, 0.5, 1, 2, 3, 4, 5, 6, or 7% or more. In some embodiments, the concentration of the surfactant in the lysis buffer is about 1% Li dodecyl sulfate. The time used for dissolution in this method may depend on the amount of surfactant used. In some embodiments, the more surfactant used, the shorter the time required for dissolution. The lysis buffer may contain a chelating agent (e.g., EDTA, EGTA).The concentration of the chelating agent in the lysis buffer may be at least about 1, 5, 10, 15, 20, 25, or 30 mM or higher. In some embodiments, the concentration of the chelating agent in the lysis buffer is about 10 mM. The lysis buffer may contain a reducing agent (e.g., β-mercaptoethanol, DTT). The concentration of the reducing agent in the lysis buffer may be at least about 1, 5, 10, 15, or 20 mM or higher. In some embodiments, the concentration of the reducing agent in the lysis buffer is about 5 mM. In some embodiments, the lysis buffer may contain about 0.1 M Tris HCl, about pH 7.5, about 0.5 M LiCl, about 1% lithium dodecyl sulfate, about 10 mM EDTA, and about 5 mM DTT.

[0087] Dissolution can be carried out at a temperature of about 4, 10, 15, 20, 25, or 30 C. Dissolution can be carried out for about 1, 5, 10, 15, or 20 minutes or longer. The lysed cells may contain at least about 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000 target nucleic acid molecules or more. Lysified cells may contain at most approximately 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000 target nucleic acid molecules or more.

[0212] In some embodiments, oligonucleotides can hybridize to nucleic acids released from cells. These nucleic acids may include DNA, such as genomic DNA, or RNA, such as messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNA containing a poly(A) tail, or any combination thereof. Nucleic acids hybridized to oligonucleotides can be used as templates for extension reactions such as reverse transcription, DNA synthesis, and other processes.

[0213] Nucleic acid molecules released from lysed cells can be associated with multiple probes on a substrate (e.g., hybridization to probes on the substrate). If the probe contains oligo-dT, the mRNA molecule can hybridize to the probe and be reverse transcribed. The oligo-dT portion of the oligonucleotide can act as a primer for the first strand synthesis of the cDNA molecule. Reverse transcription of labeled RNA molecules can be performed by adding reverse transcription primers. In some cases, the reverse transcription primers are oligo-dT primers, random hexanucleotide primers, or target-specific oligonucleotide primers. Generally, oligo-dT primers are 12-18 nucleotides long and bind to the endogenous poly(A)+ tail at the 3' end of mammalian mRNA. Random hexanucleotide primers can bind to mRNA at various complementary sites. Target-specific oligonucleotide primers typically selectively prime the target mRNA.

[0214] Reverse transcription can be repeated to produce multiple labeled cDNA molecules. The methods disclosed herein may include steps of performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 reverse transcription reactions. The methods may include steps of performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 reverse transcription reactions.

[0215] In some embodiments, nucleic acids can be removed from the substrate using chemical cleavage. For example, chemical groups or modified bases present in the nucleic acid can be used to facilitate its removal from the solid support. For example, enzymes can be used to remove nucleic acids from the substrate. For example, nucleic acids can be removed from the substrate by restriction endonuclease digestion. For example, uracil-d-glycosylase (UDG) treatment of nucleic acids containing dUTP or ddUTP can be used to remove nucleic acids from the substrate. For example, nucleic acids can be removed from the substrate using enzymes that perform nucleotide excision, such as base excision repair enzymes, such as depurine / depyrimidine (AP) endonucleases. In some embodiments, nucleic acids can be removed from the substrate using photocleavable groups and light. In some embodiments, cleavable linkers can be used to remove nucleic acids from the substrate. For example, a cleavage linker may include at least one of biotin / avidin, biotin / streptavidin, biotin / nutravidine, Ig-protein A, a photoinstability linker, an acid or base instability linker group, or an aptamer.

[0216] If the probe is gene-specific, the molecule can hybridize to the probe and be reverse transcribed and / or amplified. In some embodiments, amplification is possible after the nucleic acid has been synthesized (e.g., after reverse transcription). Amplification can be performed in a multiplexing manner, under conditions where multiple target nucleic acid sequences are amplified simultaneously. Amplification can be achieved by adding a sequencing adapter to the nucleic acid.

[0217] Amplification can be performed on a substrate, for example, using bridge amplification. A homopolymer tail can be added to the cDNA to generate ends suitable for bridge amplification using an oligo-dT probe on the substrate. In bridge amplification, the primer complementary to the 3' end of the template nucleic acid can be the first primer of each pair covalently bound to the solid particle. When a sample containing the template nucleic acid is in contact with the particle and one thermal cycle is performed, the template molecule anneals to the first primer, and the first primer extends forward by the addition of nucleotides to form a double-stranded molecule consisting of the template molecule and a newly formed DNA strand complementary to the template. In the heating step of the next cycle, the double-stranded molecule is denatured, releasing the template molecule from the particle and leaving the complementary DNA strand bound to the particle via the first primer. In the annealing step of the subsequent annealing-extension step, the complementary strand can hybridize to a second primer complementary to the segment of the complementary strand at the position removed from the first primer. This hybridization allows the complementary strand to form a bridge between the first and second primers, covalently bonded to the first primer and hybridized to the second primer. In the extension step, the second primer can be extended in the opposite direction by adding nucleotides to the same reaction mixture, thereby converting the bridge into a double-stranded bridge. The next cycle is then initiated, and the double-stranded bridge can be denatured to give two single-stranded nucleic acid molecules, each having one end bonded to the particle surface via the first and second primers, and the other end remaining unbonded. In the annealing-extension step of this second cycle, each strand can hybridize to further previously unused complementary primers on the same particle to form new single-stranded bridges. At this point, the two previously unused primers hybridized can be extended to convert the two new bridges into double-stranded bridges.

[0218] The amplification reaction may include steps to amplify at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 100% of a plurality of nucleic acids.

[0219] Amplification of labeled nucleic acids may include PCR-based or non-PCR-based methods. Amplification of labeled nucleic acids may include exponential amplification of labeled nucleic acids. Amplification of labeled nucleic acids may include linear amplification of labeled nucleic acids. Amplification can be performed by polymerase chain reaction (PCR). PCR may refer to a reaction that amplifies a specific DNA sequence in vitro by simultaneous primer extension of complementary strands of DNA. PCR may encompass derivatives of the reaction, such as, but are not limited to, RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, suppression PCR, semi-suppressive PCR, and assembly PCR.

[0220] In some cases, amplification of labeled nucleic acids involves non-PCR-based methods. Examples of non-PCR-based methods include, but are not limited to, multiple substitution amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand substitution amplification (SDA), real-time SDA, rolling circle amplification, or circle-circle amplification. Other non-PCR-based amplification methods include DNA-dependent RNA polymerase-driven RNA transcription amplification or multiple cycles of RNA-directed DNA synthesis and transcription for amplifying DNA or RNA targets, ligase chain reaction (LCR), Qβ replicase (Qβ), use of palindromic probes, strand substitution amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods in which primers are hybridized to nucleic acid sequences and the resulting double strands are cleaved before the extension reaction and amplification, strand substitution amplification using nucleic acid polymerases lacking 5' exonuclease activity, rolling circle amplification, and / or branched extension amplification (RAM).

[0221] In some cases, the methods disclosed herein further include the step of carrying out a nested polymerase chain reaction on an amplified amplicon (e.g., a target). The amplicon may be a double-stranded molecule. The double-stranded molecule may include a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized to a DNA molecule. One or both strands of the double-stranded molecule may contain a sample tag or molecular identifier label. Alternatively, the amplicon may be a single-stranded molecule. The single-stranded molecule may include DNA, RNA, or a combination thereof. The nucleic acids of the present invention may include synthetic nucleic acids or modified nucleic acids.

[0222] In some cases, the method includes a step of repeatedly amplifying a labeled nucleic acid to generate a large number of amplicons. The methods disclosed herein may include a step of performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amplification reactions. Alternatively, the method may include a step of performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 amplification reactions.

[0223] The amplification step may further include adding one or more control nucleic acids to one or more samples containing multiple nucleic acids. The amplification step may further include adding one or more control nucleic acids to multiple nucleic acids. The control nucleic acids may include a control label.

[0224] Amplification may involve the use of one or more non-natural nucleotides. Non-natural nucleotides may include photounstable and / or triggering nucleotides. Examples of non-natural nucleotides include, but are not limited to, peptide nucleic acids (PNA), morpholino nucleic acids and locked nucleic acids (LNA), as well as glycol nucleic acids (GNA) and threose nucleic acids (TNA). Non-natural nucleotides may be added to one or more cycles of the amplification reaction. The addition of non-natural nucleotides may be used to identify the product at a particular cycle or point in time of the amplification reaction.

[0225] The step of performing an amplification reaction one or more times may involve the use of one or more primers. One or more primers may contain one or more oligonucleotides. One or more oligonucleotides may contain at least about 7 to 9 nucleotides. One or more oligonucleotides may contain 12 to less than 15 nucleotides. One or more primers may anneal to at least a portion of multiple labeled nucleic acids. One or more primers may anneal to the 3' and / or 5' ends of multiple labeled nucleic acids. One or more primers may anneal to the internal regions of multiple labeled nucleic acids. The internal region may consist of at least approximately 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of multiple labeled nucleic acids. One or more primers may include a constant panel of primers. One or more primers may include at least one custom primer. One or more primers may include at least one control primer. One or more primers may include at least one housekeeping gene primer. One or more primers may include a universal primer. A universal primer may anneal to a universal primer binding site. One or more custom primers may anneal to a first sample tag, a second sample tag, a molecular identifier label, a nucleic acid, or a product thereof. One or more primers may include a universal primer and a custom primer. A custom primer may be designed to amplify one or more target nucleic acids. The target nucleic acids may include a subset of all nucleic acids in one or more samples. In some cases, the primers are probes conjugated to the array of this disclosure.

[0226] detection One or more target-specific probes can be used to detect one or more targets on a substrate. In some embodiments, the probes may be fluorescently labeled. Images of hybridized probes can be generated, for example, by fluorescence imaging. In some embodiments, the probes may be removed, and different sets of probes are used to detect different sets of targets.

[0227] The target (e.g., a molecule, amplified molecule) can be detected using, for example, a detection probe (e.g., a fluorescent probe). The array can hybridize with at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 detection probes or more. The array can hybridize with at most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 detection probes or more. In some embodiments, the array is hybridized with four detection probes.

[0228] The detection probe may contain a sequence complementary to the sequence of the target gene. The length of the detection probe may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides or more. The length of the detection probe may be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides or more. The detection probe may contain a sequence that is perfectly complementary to the sequence of the target gene (e.g., target). The detection probe may contain a sequence that is not perfectly complementary to the sequence of the target gene (e.g., target). The detection probe may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 mismatches or more sequences relative to the sequence of the target gene. The detection probe may contain at most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 mismatches or more sequences relative to the sequence of the target gene.

[0229] The detection probe may include a detectable label. Exemplary detectable labels may include fluorophores, chromophores, small molecules, nanoparticles, haptens, enzymes, antibodies, and magnetism, or any combination thereof.

[0230] Hybridized probes are imageable. The images can be used to determine the relative expression level of a target gene based on the intensity of a detectable signal (e.g., a fluorescence signal). Scanning laser fluorescence microscopes or readers can be used to obtain digital images of the emitted light from a substrate (e.g., a microarray). By scanning a focus light source (usually a laser) across the entire hybridized region, it is possible to make the hybridized substrate emit optical signals such as fluorescence. Fluorophore-specific fluorescence data can be collected and measured during the scanning operation, and then an image of the substrate can be reconstructed using appropriate algorithms, software, and computer hardware. Data is then generated by combining the expected or intended locations of probe nucleic acid features with the fluorescence intensity measured at those locations, and this is then used to determine the gene expression level or nucleic acid sequence of the target sample. The process of collecting data from expected probe locations can also be referred to as "feature extraction." Digital images can typically consist of thousands to hundreds of millions of pixels in a size range of 5 to 50 microns. Each pixel in the digital image can be represented by a 16-bit integer, allowing for 65,535 different grayscale values. The reader can sequentially capture pixels from the scanned substrate and write them to an image file that can be stored on a computer hard drive. The substrate may contain several different fluorescently tagged probe DNA samples at each spot location. The scanner excites each of the probe DNA samples by repeatedly scanning the entire substrate with a laser of the appropriate wavelength and stores them in its individual image file. The image files are then analyzed and interpreted with the help of a programmed computer.

[0231] The substrate can be imaged with a confocal laser scanner. The scanner can scan the substrate slide and generate one image for each dye used by sequentially scanning with a laser of a wavelength specific to that dye. Each dye may have a known excitation spectrum and a known emission spectrum. The scanner may include a beam splitter that reflects the laser beam toward the objective lens, thereby focusing the beam onto the surface of the slide to cause spherical fluorescence emission. A portion of the emission can travel in the reverse direction through the lens and beam splitter. After passing through the beam splitter, the fluorescence beam can be reflected by mirrors and travel through an emission filter, a focus detector lens, and a center pinhole.

[0232] Correlation between probing data and imaging data By correlating images showing the target's location with images of the sample, a spatial barcode of the target can be generated. Data from substrate scans can be correlated with images of non-dissolvable samples on the substrate. The data can be overlaid to generate a map. A map of the target's location on the sample can be constructed using information generated using the method described herein. The map can be used to determine the physical location of the target. The map can be used to identify the locations of multiple targets. Multiple targets may be of the same species, or multiple targets may be multiple different targets. For example, a map of the brain can be constructed to show the quantity and location of multiple targets.

[0233] Maps can be generated from data from a single sample. Maps can be constructed using data from multiple samples, thereby generating combined maps. Maps can be constructed with data from tens, hundreds, and / or thousands of samples. Maps composed of multiple samples can show the distribution of targets associated with regions common to multiple samples. For example, replicate assays can be displayed on the same map. At least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more replicates can be displayed (e.g., overlaid) on the same map. At most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more replicates can be displayed (e.g., overlaid) on the same map. The spatial distribution and number of targets can be represented by various statistics.

[0234] By combining data from multiple samples, it is possible to improve the positional resolution of the combined map. If individual positional measurements across the entire sample are not at least partially adjacent, the orientation of multiple samples can be aligned using shared landmarks and / or xy positions on the array. Multiplexing these methods will yield a high-resolution map of the target nucleic acid in the sample.

[0235] Data analysis and correlation can be useful in determining the presence and / or absence of specific cell types (e.g., rare cells, cancer cells). Data correlation can be useful in determining the relative ratio of target nucleic acids at identifiable locations, either intracellular or within a sample.

[0236] The methods and compositions disclosed herein may be companion diagnostics for medical professionals (e.g., pathologists). In this case, a subject can be diagnosed by visually observing pathological images and correlating the images with gene expression (e.g., by identifying oncogene expression). These methods and compositions may be useful for identifying cells in a cell population and determining the genetic heterogeneity of cells in a sample. These methods and compositions may be useful for determining the genotype of a sample.

[0237] This disclosure provides a method for producing replicates of a substrate. The substrate can be reprobed with different probes for different target genes or selectively selected for specific genes. For example, a sample can be placed on a substrate containing multiple oligo(dT) probes. mRNA can hybridize to the probes. A replicate substrate containing oligo(dT) probes can be brought into contact with an initial slide to produce mRNA replicates. A replicate substrate containing RNA gene-specific probes can be brought into contact with an initial slide to produce replicates.

[0238] mRNA can be reverse transcribed into cDNA. cDNA can be homopolymerized and / or amplified (e.g., by bridge amplification). The array can be in contact with the replicate array. The replicate array may contain gene-specific probes that can bind to the target cDNA. The replicate array may contain poly(A) probes that can bind to cDNA having a polyadenylated sequence.

[0239] The number of replicas that can be produced may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more. The number of replicas that can be produced may be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more.

[0240] In some embodiments, the initial substrate includes multiple gene-specific probes, and the replicated substrate includes the same gene-specific probe, or different probes corresponding to the same gene as the gene-specific probe.

[0241] Imaging Samples brought into contact with a substrate can be analyzed (e.g., using immunohistochemistry, staining, and / or imaging). An exemplary immunohistochemical method may involve reacting a labeled probe biomaterial, obtained by introducing a label to a substance recognizable to the target biomaterial, with a tissue section in order to visualize the target biomaterial present on the tissue section through a specific binding reaction between biomaterials.

[0242] In tissue diagnostic specimens, the tissue sample can be fixed with a suitable fixative, typically formalin, and embedded in molten paraffin wax. The wax block can be cut with a microtome to produce thin slices of paraffin containing the tissue. The specimen slices can be applied to a substrate, air-dried, and then heated to adhere the specimen to a glass slide. Residual paraffin can be dissolved with a suitable solvent, typically xylene or toluene. These so-called deparaffinizing solvents can be removed with a wash-dehydrating reagent before staining. Slices can be prepared from frozen specimens, briefly fixed in 10% formalin, and then injected with a dehydrating reagent. The dehydrating reagent can be removed before staining with aqueous stains.

[0243] In some embodiments, Papanicolaou staining techniques can be used (e.g., progressive staining and / or hematoxylin-eosin [H&E], i.e., degenerative staining). HE (hematoxylin-eosin) staining uses hematoxylin and eosin as dyes. Hematoxylin is a blue-violet dye that stains basophilic tissues such as cell nuclei, bone tissue, parts of cartilage tissue, and serous components. Eosin is a red to pink dye that stains eosinophilic tissues such as cytoplasm, connective tissue of soft tissues, red blood cells, fibrin, and endocrine granules.

[0244] Immunohistochemistry (IHC) can also be referred to as "immunological staining" because it uses a color-developing process to visualize antigen-antibody reactions that cannot be recognized by other methods (hereafter, the term "immunohistochemical staining" can be used for immunohistochemistry). Lectin staining is a technique that uses lectins to detect glycans in tissue samples by utilizing the properties of lectins that bind to specific glycans in a non-immunological and specific way.

[0245] HE staining, immunohistochemistry, and lectin staining can be used, for example, to detect the location of cancer cells in a cell sample. For example, if it is desirable to confirm the location of cancer cells in a cell sample, a pathologist can prepare tissue sections and place them on the substrates of this disclosure to determine the presence or absence of cancer cells in the cell sample. The sections on the array can be subjected to HE staining, imaging, or any immunohistochemical analysis to obtain their morphological information and / or any other identifying features (e.g., the presence or absence of rare cells). The sample is soluble, and the presence or absence of nucleic acid molecules can be determined using the methods of this disclosure. Nucleic acid information can be compared with images (e.g., spatial comparison), thereby indicating the spatial location of nucleic acids in the sample.

[0246] In some embodiments, the tissue is stained with a staining enhancer (e.g., a chemiotommulin enhancer). Examples of histochemical permeability enhancers that facilitate the penetration of stains into tissues include, but are not limited to, polyethylene glycol (PEG), surfactants such as polyoxyethylene sorbitan, polyoxyethylene ethers (polyoxyethylene sorbitan monolaurate (Tween 20) and other Tween derivatives), polyoxyethylene 23-lauryl ether (Brij 35), Triton X-100, Brij, Nonidet P-40, surfactant-like substances such as lysolecithin, saponins, nonionic surfactants such as TRITON® X-100, aprotic solvents such as dimethyl sulfoxide (DMSO), ethers such as tetrahydrofuran, dioxane, esters such as ethyl acetate, butyl acetate, isopropyl acetate, hydrocarbons such as toluene, chlorinated solvents such as dichloromethane, dichloroethane, chlorobenzene, ketones such as acetone, nitriles such as acetonitrile, and / or other agents that increase cell membrane permeability.

[0247] In some embodiments, compositions are provided that facilitate the staining of mammalian tissue samples. The compositions may comprise a staining agent, for example, hematoxylin, or hematoxylin and eosin-Y, at least one histochemical penetration enhancer, for example, a surfactant, an aprotic solvent, and / or PEG, or any combination thereof.

[0248] In some embodiments, the sample is imaged (e.g., before or after IHC, or without IHC). Imaging may include microscopy, e.g., light-area imaging, gradient illumination, dark-area imaging, dispersion staining, phase contrast, differential interference contrast, interference reflection microscopy, fluorescence, confocal, electron microscopy, transmission electron microscopy, scanning electron microscopy, single-plane illumination, or any combination thereof. Imaging may include the use of negative stains (e.g., nigrosine, ammonium molybdate, uranyl acetate, uranyl formate, phosphotungstic acid, osmium tetroxide). Imaging may include the use of heavy metals capable of scattering electrons (e.g., gold, osmium).

[0249] Imaging may include imaging of a portion of the sample (e.g., slides / arrays). Imaging may include imaging of at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% of the sample. Imaging may include imaging of at most 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% of the sample. Imaging can be performed in separate steps (e.g., images do not necessarily need to be adjacent). Imaging may include the acquisition of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more different images. Imaging may include the acquisition of at most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more different images.

[0250] Figure 9 shows an exemplary embodiment of the homopolymer tailing method of the present disclosure. The present disclosure provides a substrate 910 comprising a plurality of probes 905 bonded to the surface of the substrate. The substrate 910 may be a microarray. The plurality of probes 905 may comprise oligo(dT). The plurality of probes 905 may comprise gene-specific sequences. The plurality of probes 905 may comprise probabilistic barcodes. A sample (e.g., cells) 915 can be placed and / or grown on the substrate 910. The substrate containing the sample can be analyzed, for example, by imaging and / or immunohistochemistry 920. The sample 915 can be lysed on the substrate 910 925. The nucleic acid 930 of the sample 915 can be associated (e.g., hybridized) with the plurality of probes 905 on the substrate 910. In some embodiments, the nucleic acid 930 can be reverse transcribed, homopolymer tailed, and / or amplified (e.g., bridge amplification). The amplified nucleic acid can be examined 935 using a detection probe 940 (e.g., a fluorescent probe). The detection probe 940 may be a gene-specific probe. The binding site of the detection probe 940 on the substrate 910 can be correlated with an image of the substrate to generate a map showing the spatial location of nucleic acids in the sample.

[0251] In some embodiments, the method may include the preparation 945 of a replica 946 of the original substrate 910. The replica substrate 946 may contain a plurality of probes 931. The plurality of probes 931 may be identical to the plurality of probes 905 on the original substrate 910. The plurality of probes 931 may be different from the plurality of probes 905 on the original substrate 910. For example, the plurality of probes 905 may be oligo(dT) probes, while the plurality of probes 931 on the replica substrate 946 may be gene-specific probes. The replica substrate can be processed in the same way as the original substrate, such as by examination with detection (e.g., fluorescent) probes.

[0252] sample cell The samples used in the methods of this disclosure may include one or more cells. A sample may mean one or more cells. In some embodiments, the cells are cancer cells extracted from cancerous tissue, such as breast cancer, lung cancer, colon cancer, prostate cancer, ovarian cancer, pancreatic cancer, brain cancer, melanoma, and non-melanoma skin cancer. In some cases, the cells are derived from cancer but collected from bodily fluids (e.g., circulating tumor cells). Examples of cancer, but not limited to, include adenoma, adenocarcinoma, squamous cell carcinoma, basal cell carcinoma, small cell carcinoma, large cell anaplastic carcinoma, chondrosarcoma, and fibrosarcoma.

[0253] In some embodiments, the cells are those infected with a virus and containing viral oligonucleotides. In some embodiments, the viral infection may be caused by a virus selected from the group consisting of double-stranded DNA viruses (e.g., adenovirus, herpesvirus, poxvirus), single-stranded (+strand or "sense") DNA viruses (e.g., parvovirus), double-stranded RNA viruses (e.g., reovirus), single-stranded (+strand or sense) RNA viruses (e.g., picornavirus, togavirus), single-stranded (+strand or antisense) RNA viruses (e.g., orthomyxovirus, rhabdovirus), single-stranded (+strand or sense) RNA-RT viruses (e.g., retroviruses), and double-stranded DNA-RT viruses (e.g., hepadnavirus). Exemplary viruses include, but are not limited to, SARS, HIV, coronavirus, Ebola, malaria, dengue fever, hepatitis C, hepatitis B, and influenza.

[0254] In some embodiments, the cells are bacteria. These may include either Gram-positive or Gram-negative bacteria. Examples of bacteria that can be analyzed using the methods, devices, and systems of this disclosure include, but are not limited to, the genera Actinomedurae, Actinomyces israelii, Bacillus anthracis, Bacillus cereus, Clostridium botulinum, Clostridium difficile, Clostridium perfringens, Clostridium tetani, Corynebacterium, Enterococcus faecalis, and Listeria monocytogenes. Examples include monocytogenes, the genus Nocardia, Propionibacterium acnes, Staphylococcus aureus, Staphylococcus epiderm, Streptococcus mutans, and Streptococcus pneumoniae. Gram-negative bacteria are not limited to this group, but include Afipia felis, Bacteriodes, Bartonella bacilliformis, Bortadella pertussis, Borrelia burgdorferi, Borrelia recurrentis, Brucella, and Calymmatobacterium granulomatis.Campylobacter granulomatis, Escherichia coli, Francisella tularensis, Gardnerella vaginalis, Haemophilius aegyptius, Haemophilius ducreyi, Haemophilius influenziae, Heliobacter pylori, Legionella pneumophila, Leptospira interrogans, Neisseria meningitidia, Porphyromonas gingivalis Examples include *Gingivalis*, *Providencia sturti*, *Pseudomonas aeruginosa*, *Salmonella enteridis*, *Salmonella typhi*, *Serratia marcescens*, *Shigella boydii*, *Streptobacillus moniliformis*, *Streptococcus pyogenes*, *Treponema pallidum*, *Vibrio cholerae*, *Yersinia enterocolitica*, and *Yersinia pestis*. Other bacteria include Myobacterium avium, Myobacterium leprae, Myobacterium tuberculosis, and Bartonella henselae.Chlamydia psittaci, Chlamydia trachomatis, Coxiella burnetii, Mycoplasma pneumoniae, Rickettsia akari, Rickettsia prowazekii, Rickettsia rickettsii, Rickettsia tsutsugamushi, Rickettsia typhi, Ureaplasma urealyticum, Diplococcus pneumoniae, Ehrlichia shaffinsis Examples include *Chephrococcus chafensis*, *Enterococcus faecium*, and *Meningococcus*.

[0255] In some embodiments, the cells are fungi. Examples of fungi that can be analyzed using the methods, devices, and systems of this disclosure include, but are not limited to, the genera Aspergillus, Candidae, Candida albicans, Coccidioides immitis, Cryptococci, and combinations thereof.

[0256] In some embodiments, the cells are protozoa or other parasites. Examples of parasites analyzed using the methods, devices, and systems of this disclosure include, but are not limited to, Balantidium coli, Cryptosporidium parvum, Cyclospora cayatanensis, Encephalitozoa, Entamoeba histolytica, Enterocytozoon bieneusi, Giardia lamblia, Leishmaniae, Plasmodium, Toxoplasma gondii, Trypanosomae, and trapezoidal amoeba. Examples include amoeba, worms (e.g., helminths), and especially parasites, such as, but not limited to, Nematoda (roundworms, e.g., whipworms, hookworms, pinworms, filamentous worms, etc.) and Cestoda (e.g., tapeworms).

[0257] As used herein, the term “cell” may mean one or more cells. In some embodiments, cells are normal cells, for example, human cells at various developmental stages or human cells of various organ or tissue types (e.g., leukocytes, erythrocytes, platelets, epithelial cells, endothelial cells, neurons, glial cells, fibroblasts, skeletal muscle cells, smooth muscle cells, gametes, or cells of the heart, lungs, brain, liver, kidneys, spleen, pancreas, thymus, bladder, stomach, colon, or small intestine). In some embodiments, cells may be undifferentiated human stem cells or differentiated human stem cells. In some embodiments, cells may be fetal human cells. Fetal human cells may be obtained from a pregnant mother with a fetus. In some embodiments, cells are rare cells. Rare cells may include, for example, circulating tumor cells (CTCs), circulating epithelial cells, circulating endothelial cells, circulating endometrial cells, circulating stem cells, stem cells, undifferentiated stem cells, cancer stem cells, bone marrow cells, progenitor cells, foam bubbles, mesenchymal cells, trophoblasts, immune system cells (host or graft), cell fragments, cellular organelles (e.g., mitochondria or nucleus), and pathogen-infected cells.

[0258] In some embodiments, the cells are non-human cells (e.g., other types of mammalian cells) (e.g., mouse, rat, pig, dog, cattle, or horse). In some embodiments, the cells are other types of animal or plant cells. In other embodiments, the cells may be any prokaryotic or eukaryotic cells.

[0259] In some embodiments, the first cell sample is obtained from a person without the disease or condition, and the second cell sample is obtained from a person with the disease or condition. In some embodiments, the persons are different. In some embodiments, the persons are the same, but the cell samples are obtained at different points in time. In some embodiments, the person is a patient, and the cell sample is a patient sample. The disease or condition may be cancer, bacterial infection, viral infection, inflammatory disease, neurodegenerative disease, fungal disease, parasitic disease, genetic disorder, or any combination thereof.

[0260] In some embodiments, cells suitable for use in the methods of this disclosure are in a size range of about 2 micrometers to about 100 micrometers in diameter. In some embodiments, cells have a diameter of at least 2 micrometers, at least 5 micrometers, at least 10 micrometers, at least 15 micrometers, at least 20 micrometers, at least 30 micrometers, at least 40 micrometers, at least 50 micrometers, at least 60 micrometers, at least 70 micrometers, at least 80 micrometers, at least 90 micrometers, or at least 100 micrometers. In some embodiments, cells have a diameter of at most 100 micrometers, at most 90 micrometers, at most 80 micrometers, at most 70 micrometers, at most 60 micrometers, at most 50 micrometers, at most 40 micrometers, at most 30 micrometers, at most 20 micrometers, at most 15 micrometers, at most 10 micrometers, at most 5 micrometers, or at most 2 micrometers. Cells may have a diameter of any value in the range of about 5 micrometers to about 85 micrometers, for example. In some embodiments, the cells have a diameter of approximately 10 micrometers.

[0261] In some embodiments, cells are sorted before associating them with beads and / or microwells. For example, cells can be sorted by fluorescence-activated cell sorting, magnetoactivated cell sorting, or, for example, flow cytometry. Cells can be filtered by size. In some cases, the retainate contains cells to be associated with beads. In some cases, the flow-through contains cells to be associated with beads.

[0262] A sample can be nucleic acid. A sample can contain nucleic acid. A sample can be a single cell. If a sample is a single cell, the cell can be lysed to release nucleic acid from the single cell. A sample can be multiple cells. Samples in different wells can originate from the same subject. Samples in different wells can originate from different subjects. Samples in different wells can originate from different tissues (e.g., brain, heart, lungs, kidneys, spleen). Samples in different wells can originate from the same tissue.

[0263] Diffusion across the substrate When a sample (e.g., cells) is labeled with a probabilistic barcode and / or combinatorial barcode according to the method of this disclosure, the cells may be lysed. Lysis of cells may result in diffusion of the lysed contents (e.g., cellular contents) away from the initial location of lysis. In other words, the lysed contents can move to a surface area larger than the surface area occupied by the cells.

[0264] The diffusion of a sample solution mixture (e.g., containing a target) can be controlled by various parameters, including, but not limited to, the viscosity of the solution mixture, the temperature of the solution mixture, the size of the target, the size of the physical barrier in the substrate, and the concentration of the solution mixture. For example, the temperature of the dissolution reaction can be set to at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, or 40C or higher. The viscosity of the solution mixture can be modified, for example, by adding a thickening agent (e.g., glycerol, beads) to slow the diffusion rate. The viscosity of the solution mixture can be modified, for example, by adding a diluent (e.g., water) to increase the diffusion rate. The substrate may include physical barriers (e.g., wells, microwells, microhills) that can modify the diffusion rate of the sample target. The concentration of the dissolution mixture can be modified to increase or decrease the diffusion rate of the sample target. The concentration of the dissolution mixture can be increased or decreased by at least 1, 2, 3, 4, 5, 6, 7, 8, or 9 or more times. The concentration of the dissolution mixture can be increased or decreased by at most 1, 2, 3, 4, 5, 6, 7, 8, or 9 or more times.

[0265] The diffusion rate can be increased. The diffusion rate can be decreased. The diffusion rate of the dissolved mixture can be increased or decreased by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more times compared to the unmodified dissolved mixture. The diffusion rate of the dissolved mixture can be increased or decreased by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more times compared to the unmodified dissolved mixture. The diffusion rate of the dissolved mixture can be increased or decreased by at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% compared to the unmodified dissolved mixture. The diffusion rate of the dissolved mixture can be increased or decreased by at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% compared to the unmodified dissolved mixture.

[0266] Data analysis and display software Data analysis and visualization of the target's spatial resolution. This disclosure provides a method for estimating the number and location of targets using probabilistic barcoding and / or combinatorial barcoding and digital counting. The data obtained from the methods of this disclosure can be visualized on a map. The data obtained from the methods of this disclosure can be visualized on a map of a microwell array (for example, so that the results for each sample can be traced to their location in the microwell array). A map of the number and location of targets in a sample can be constructed using information generated using the methods described herein. The map can be used to determine the physical location of the targets. The map can be used to identify the locations of multiple targets. The multiple targets may be of the same species, or the multiple targets may be multiple different targets. For example, it is possible to construct a map of the brain to show the digital count and location of multiple targets.

[0267] Maps can be generated from data from a single sample. Maps can be constructed using data from multiple samples, thereby generating combined maps. Maps can be constructed using data from tens, hundreds, and / or thousands of samples. Maps composed of multiple samples can show the distribution of digital counts of targets associated with regions common to multiple samples. For example, replicate assays can be displayed on the same map. At least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 replicates or more can be displayed (e.g., overlaid) on the same map. At most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 replicates or more can be displayed (e.g., overlaid) on the same map. The spatial distribution and number of targets can be represented by various statistics.

[0268] In some embodiments of the instrument system, the system may include a computer-readable medium containing code for providing data analysis of sequence datasets generated by performing a single-cell stochastic barcoding assay. Examples of data analysis functions that may be provided by data analysis software include, but are not limited to, (i) algorithms for decoding / demultiplexing sample labels, cell labels, molecular labels, and target sequence data provided by sequencing a probabilistic barcode and / or combinatorial barcode reagent library generated during assay execution; (ii) algorithms for determining read count / gene / cell and unique transcript molecule count / gene / cell; (iii) statistical analysis of sequence data for clustering cells by gene expression data or predicting confidence intervals for determinations such as transcript molecule count / gene / cell; (iv) algorithms for identifying rare cell subpopulations using, for example, principal component analysis, hierarchical clustering, k-mean clustering, self-organizing maps, neural networks; (v) sequence alignment functions for aligning gene sequence data to known reference sequences and for detecting mutations, polymorphism markers, and splice variants; and (vi) automated clustering of molecular labels to compensate for amplification or sequencing errors. In some embodiments, commercially available software may be used to perform all or part of the data analysis. For example, Seven Bridges (https: / / www.sbgenomics.com / ) software can be used to compile a table showing the number of copies of one or more genes present in each cell of a whole cell collection. In some embodiments, the data analysis software may include an option to output sequencing results in a useful graphical format, such as a heatmap showing the number of copies of one or more genes present in each cell of a cell population.In some embodiments, the data analysis software may further include algorithms for extracting biological meaning from sequencing results, for example, by correlating the number of copies of one or more genes present in each cell of a cell population with cells of a certain type, a certain type of rare cell type, or cells derived from subjects with a specific disease or condition. In some embodiments, the data analysis software may further include algorithms for comparing cell populations across different biological samples.

[0269] In some embodiments, all data analysis functions can be packaged within a single software package. In some embodiments, a complete set of data analysis capabilities may comprise a single software package. In some embodiments, the data analysis software may be a standalone package available to the user independently of the assay instrument system. In some embodiments, the software may be web-based and may allow users to share data.

[0270] System processor and network Generally, the computer or processor included in the equipment system of this disclosure can be further understood as a logical device capable of reading instructions from a medium or network port that can be optionally connected to a server having a fixed medium. The system may include a CPU, disk drives, optional input devices such as a keyboard and mouse, and an optional monitor. Data communication is achievable to a server at a local or remote location via a specified communication medium. The communication medium may include any means of sending and receiving data. For example, the communication medium may be a network connection, a wireless connection, or an Internet connection. Such a connection can provide communication via the World Wide Web. Data relating to this disclosure can be transmitted via such network or connection for reception or viewing by a party.

[0271] Exemplary embodiments of the first architectural example of a computer system are available for use in connection with the embodiments of the present disclosure. The computer system example may include a processor for processing instructions. Examples of processors, but are not limited to, Intel Xeon® processors, AMD Opteron® processors, Samsung 32-bit RISC ARM 1176JZ(F)-S v1.0® processors, ARM Cortex-A8 Samsung S5PC100® processors, ARM Cortex-A8 AppleA4® processors, Marvell PXA 930® processors, or functionally equivalent processors. Multithreading of execution is available for parallel processing. In some embodiments, a processor with multiple processors or multiple cores is also available, whether it is a clustered single computer system or a networked distributed system including multiple computers, mobile phones, or personal digital assistant devices.

[0272] A high-speed cache can be attached to or introduced into the processor to provide high-speed memory for instructions or data recently or frequently used by the processor. The processor can be connected to the northbridge via the processor bus. The northbridge is connected to random access memory (RAM) via the memory bus, and the processor manages access to RAM. The northbridge can also be connected to the southbridge via the chipset bus. The southbridge, in turn, is connected to the peripheral bus. The peripheral bus can be, for example, PCI, PCI-X, PCI Express, or other peripheral buses. The northbridge and southbridge are often referred to as the processor chipset and manage data transfer between the processor, RAM, and peripheral elements on the peripheral bus. In some alternative architectures, the functionality of the northbridge can be integrated into the processor instead of using a separate northbridge chip.

[0273] The system may include an accelerator card coupled to a bus for peripherals. The accelerator may include a field-programmable gate array (FPGA) or other hardware to accelerate certain processes. For example, an accelerator could be used for adaptive data restructuring or to evaluate algebraic expressions used in extended set processing.

[0274] The software and data are retrievable in external storage devices and can be loaded into RAM or cache for use by the processor. The system includes an operating system for management system resources. Examples of operating systems include, but are not limited to, Linux®, Windows®, macOS®, BlackBerry OS®, iOS®, and other functionally equivalent operating systems, as well as application software running on an operating system for managing data storage and optimization according to embodiments of the present invention.

[0275] In this example, the system also includes a network interface card (NIC) connected to a peripheral bus that provides a network interface to external storage devices such as network-attached storage (NAS) and other computer systems that can be used for distributed parallel processing.

[0276] An exemplary diagram of the network may include multiple computer systems, multiple mobile phones and personal personal information terminals, and network-attached storage (NAS). In this embodiment, the system can manage data storage and optimize data access to data stored in the network-attached storage (NAS). Mathematical models can be used on the data and evaluated using distributed parallel processing computer systems and mobile phones and personal personal information terminal systems. The computer systems and mobile phones and personal personal information terminal systems can also provide parallel processing for adaptive data restructuring of data stored in the network-attached storage (NAS). A wide variety of other computer architectures and systems can be used in connection with various embodiments of the present invention. For example, blade servers can be used to provide parallel processing. Processor blades can be connected via a backplane to provide parallel processing. Storage can also be connected to a backplane or exist as network-attached storage (NAS) via a separate network interface.

[0277] In some embodiments, the processor can maintain its own memory space and transmit data via a network interface to the backplane or to other connectors for parallel processing by other processors. In other embodiments, some or all of the processors can use a shared virtual address memory space.

[0278] An exemplary block diagram of a multiprocessor computer system may include a shared virtual address memory space according to an embodiment. The system may include multiple processors that can access the shared memory subsystem. The system may incorporate multiple programmable hardware memory algorithm processors (MAPs) in the memory subsystem. Each MAP may include memory and one or more field-programmable gate arrays (FPGAs). A MAP can provide configurable functional units, and a particular algorithm or part of an algorithm can be provided to the FPGA for processing in close cooperation with its respective processor. For example, a MAP can be used to evaluate algebraic expressions relating to a data model and, in an embodiment, to perform adaptive data restructuring. In this example, each MAP is globally accessible by all processors for these purposes. In one configuration, each MAP can use direct memory access (DMA) to access its associated memory, thereby enabling it to perform tasks asynchronously and independently of its respective microprocessor. In this configuration, a MAP can directly feed results to other MAPs for pipelined and parallel execution of algorithms.

[0279] The computer architectures and systems described above are merely examples, and a wide variety of other computer, mobile phone, and personal information terminal architectures and systems can be used in relation to the examples, including systems using general-purpose processors, co-processors, FPGAs, and other programmable logic devices, systems-on-a-chip (SOCs), application-specific integrated circuits (ASICs), and any combination of other processing and logic elements. In some embodiments, all or part of the computer system can be implemented in software or hardware. Any variety of data storage media, including random-access memory, hard drives, flash memory, tape drives, disk arrays, network-attached storage (NAS), and other local or distributed data storage devices and systems, can be used in relation to the examples.

[0280] In some embodiments, the computer subsystems of the present disclosure can be implemented using software modules running on any of the above or other computer architectures and systems. In other embodiments, the functionality of the system can be partially or fully implemented by firmware, programmable logic devices, such as field-programmable gate arrays (FPGAs), system-on-chip (SOCs), application-specific integrated circuits (ASICs), or other processing and logic elements. For example, set processors and optimizers can be implemented with hardware acceleration using hardware accelerator cards, such as accelerator cards.

[0281] kit Disclosed herein are kits for performing single-cell probabilistic barcoding and / or combinatorial barcoding assays. A kit may comprise any composition or mixture of compositions of the disclosure (e.g., any type and any number of any combinatorial barcoding reagents). A kit may comprise one or more combinatorial barcoding reagents of the disclosure. A kit may comprise one or more probabilistic barcoding reagents of the disclosure. A kit may comprise one or more substrates (e.g., microwell arrays) packaged as a self-supporting substrate (or chip) containing one or more microwell arrays, or in one or more flow cells or cartridges. One or more substrates of a kit may comprise combinatorial barcoding reagents and / or probabilistic barcoding reagents preloaded in the wells. In some cases, the probabilistic barcoding reagents may be preloaded in the wells of the substrate, and the kit may comprise combinatorial barcoding reagents for user addition to the wells.

[0282] In some cases, the kit may include a set of combinatorial barcoding reagents. The second combinatorial barcoding reagent may include one combinatorial barcoding reagent containing a target-specific region (e.g., to either the sense strand or the antisense strand), while the other combinatorial barcoding reagent may not contain a target-specific region but may include a linker that binds all combinatorial barcoding reagents together (e.g., so that the combinatorial barcoding reagents bind together to one end of the target nucleic acid (see Figure 2C)).

[0283] The kit may comprise one or more solid carrier suspension agents, wherein the individual solid carriers within the suspension comprise a plurality of bound probability barcodes of the present disclosure. The kit may comprise probability barcodes that are not bound to the solid carriers. The kit may comprise a sealant. In some embodiments, the kit may further comprise mechanical fasteners for mounting a self-supporting substrate to form reaction wells that facilitate pipetting of samples and reagents into the substrate.

[0284] The kit may further include reagents for performing probabilistic barcoding assays, such as lysis buffer, rinse buffer, or hybridization buffer. The kit may further include reagents (e.g., enzymes, primers, dNTPs, NTPs, RNase inhibitors, or buffers) for performing nucleic acid elongation reactions, such as reverse transcription or primer elongation. The kit may further include reagents (e.g., enzymes, universal primers, sequencing primers, target-specific primers, or buffers) for preparing sequencing libraries by amplification reactions. The kit may include reagents for homopolymer tailing of molecules (e.g., terminal transferase enzymes and dNTPs). The kit may include reagents for performing any enzymatic cleavage of the present disclosure (e.g., ExoI nuclease restriction enzymes). Reagents can be preloaded into the wells of the substrate (e.g., combinatorial barcoding reagents).

[0285] The kit may include sequencing library amplification primers. The kit may include the second strand synthesis primers of this disclosure. The kit may include any primers of this disclosure (e.g., gene-specific primers, random multimers, sequencing primers, and universal primers).

[0286] The kit may include one or more molds, for example, a mold containing an array of micropillars for casting a substrate (e.g., a microwell array) and one or more solid carriers (e.g., beads), provided that the individual beads in the suspension contain multiple combined probability barcodes of the present disclosure. The kit may further include materials for use when casting a substrate (e.g., agarose, hydrogel, PDMS, optical adhesive, etc.).

[0287] The kit may comprise one or more substrates preloaded with a solid carrier containing multiple bounded probability barcodes as disclosed herein. In some cases, there may be one solid carrier per microwell of the substrate. In some embodiments, the multiple probability barcodes may be bound directly to the surface of the substrate rather than to the solid carrier. In any of these embodiments, one or more microwell arrays may be provided in the form of self-supporting substrates (or chips) or filled into flow cells or cartridges.

[0288] In some embodiments, the kit may comprise one or more cartridges incorporating one or more substrates. In some embodiments, one or more cartridges further comprise one or more preloaded solid carriers, wherein the individual solid carriers in the suspension comprise one of the plurality of bound probabilistic barcodes of the present disclosure. In some embodiments, beads are pre-dispersed in one or more microwell arrays of the cartridge. In some embodiments, the beads are preloaded and storable in the form of a suspension in the reagent wells of the cartridge. In some embodiments, one or more cartridges further comprise other assay reagents preloaded and stored in the reagent reservoir of the cartridge.

[0289] A kit may generally include instructions for performing one or more of the methods described herein. Instructions included in a kit may be attached to the packaging material or included as accompanying documentation. Instructions are typically, but not limited to, documents or printed materials. Any medium capable of storing and communicating such instructions to an end user is contemplated in this disclosure. Such media include, but are not limited to, electronic storage media (e.g., magnetic disks, tapes, cartridges, chips), optical media (e.g., CD-ROMs), and RF tags. As used herein, the term “instructions” may include the address of an internet site providing the instructions.

[0290] device Flow Cell Microwell array substrates can be packaged within a flow cell that provides a convenient interface with the rest of the fluid handling system and facilitates the exchange of fluids, such as cells and solid carrier suspensions, lysis buffers, and rinse buffers delivered to the microwell array and / or emulsion droplets. Design features may include (i) one or more inlet ports for introducing cell samples, solid carrier suspensions, or other assay reagents; (ii) one or more microwell array chambers designed to provide uniform filling and efficient fluid exchange while minimizing back-eddy or dead zones; and (iii) one or more outlet ports for delivering fluids to a sample collection point or waste reservoir. Flow cell designs may include multiple microarray chambers interface with multiple microwell arrays so that one or more different cell samples are processed in parallel. The design of the flow cell may further include features that generate a uniform flow rate profile, i.e., a “plug flow” across the entire width of the array chamber to provide more uniform delivery of cells and beads to the microwells, for example, by using a porous barrier located near the chamber inlet and upstream of the microwell array as a “flow diffuser,” or by dividing each array chamber into several subsections that cover the entire array area as a whole but through which divided inlet fluid streams flow in parallel. In some embodiments, the flow cell may house or introduce two or more microwell array substrates. In some embodiments, an integrated microwell array / flow cell assembly may constitute a fixed component of the system. In some embodiments, the microwell array / flow cell assembly may be removable from the device.

[0291] Generally, in flow cell designs, the dimensions of the fluid channel and array chamber will be optimized to (i) provide uniform delivery of cells and beads to the microwell array, and (ii) minimize sample and reagent consumption. In some embodiments, the width of the fluid channel may be 50 μm to 20 mm. In other embodiments, the width of the fluid channel may be at least 50 μm, at least 100 μm, at least 200 μm, at least 300 μm, at least 400 μm, at least 500 μm, at least 750 μm, at least 1 mm, at least 2.5 mm, at least 5 mm, at least 10 mm, at least 20 mm, at least 50 mm, at least 100 mm, or at least 150 mm. In yet another embodiment, the width of the fluid channel may be at most 150 mm, at most 100 mm, at most 50 mm, at most 20 mm, at most 10 mm, at most 5 mm, at most 2.5 mm, at most 1 mm, at most 750 μm, at most 500 μm, at most 400 μm, at most 300 μm, at most 200 μm, at most 100 μm, or at most 50 μm. In one embodiment, the width of the fluid channel is approximately 2 mm. The width of the fluid channel may fall within any range bounded by any of these values ​​(for example, approximately 250 μm to approximately 3 mm).

[0292] In some embodiments, the depth of the fluid channel may be 50 μm to 2 mm. In other embodiments, the depth of the fluid channel may be at least 50 μm, at least 100 μm, at least 200 μm, at least 300 μm, at least 400 μm, at least 500 μm, at least 750 μm, at least 1 mm, at least 1.25 mm, at least 1.5 mm, at least 1.75 mm, or at least 2 mm. In yet another embodiment, the depth of the fluid channel may be at most 2 mm, at most 1.75 mm, at most 1.5 mm, at most 1.25 mm, at most 1 mm, at most 750 μm, at most 500 μm, at most 400 μm, at most 300 μm, at most 200 μm, at most 100 μm, or at most 50 μm. In one embodiment, the depth of the fluid channel is approximately 1 mm. The depth of the fluid channel can fall within any range bounded by any of these values ​​(for example, approximately 800 μm to approximately 1 mm).

[0293] Flow cells can be fabricated using various techniques and materials known to those skilled in the art. For example, flow cells can be fabricated as individual components and subsequently mechanically clamped or permanently bonded to a microwell array substrate. Examples of preferred fabrication techniques include conventional machining, CNC machining, injection molding, 3D printing, alignment and lamination of one or more layers of laser-cut or die-cut polymer films, or several microfabrication techniques, such as photolithography and wet chemical etching, dry etching, deep reactive ion etching, or laser micromachining. After fabricating the flow cell portion, it can be mechanically bonded to the microwell array substrate, for example, by clamping it to the microwell array substrate (with or without a gasket), or it can be directly bonded to the microwell array substrate using any of the various techniques known to those skilled in the art (depending on the choice of materials used), such as anodic bonding, thermal bonding, or using any of the various adhesives or adhesive films, such as epoxy, acrylic, silicone, UV-curable, polyurethane, or cyanoacrylate adhesives.

[0294] Flow cells can be fabricated using a variety of materials known to those skilled in the art. Generally, the choice of material depends on the choice of fabrication technique, and vice versa. Suitable materials include, but are not limited to, silicon, fused silica, glass, various polymers such as polydimethylsiloxane (PDMS, elastomer), polymethyl methacrylate (PMMA), polycarbonate (PC), polypropylene (PP), polyethylene (PE), high-density polyethylene (HDPE), polyimide, cyclic olefin polymers (COP), cyclic olefin copolymers (COC), polyethylene terephthalate (PET), epoxy resins, metals (e.g., aluminum, stainless steel, copper, nickel, chromium, and titanium), non-adherent materials such as Teflon® (PTFE), or combinations thereof.

[0295] cartridge In some embodiments of the system, flow cell-bound or unbound microwell arrays may be packaged in a consumerable cartridge that interfaces with the instrument system. Cartridge design features include: (i) one or more inlet ports for forming a fluid connection with the instrument or for manually introducing cell samples, bead suspensions, or other assay reagents into the cartridge; (ii) one or more bypass channels, i.e., for self-metering of cell samples and bead suspensions to avoid overfilling or backflow; (iii) one or more chambers in which one or more integrated microwell array / flow cell assemblies or microarray substrates are placed; (iv) an integrated miniature pump or other fluid drive mechanism for controlling the fluid flow through the device; and (v) pre-loaded reagents (e.g., bead suspensions, combinatorial barcode reagents, etc.). (vi) one or more vents for providing a passage for trapped air to escape, (vii) one or more sample and reagent waste reservoirs, (viii) one or more outlet ports for forming a fluid connection with the instrument or providing a processing sample collection point, (ix) a mechanical interface feature for reproducibly positioning a removable consumerable cartridge relative to the instrument system and providing access to bring the external magnet and the microwell array into very close proximity, (x) an integrated temperature control component or thermal interface for providing good thermal contact with the instrument system, and (xi) an optical interface feature for use in optical inspection of the microwell array, such as a transparent window.

[0296] The cartridge can be designed to process one or more samples in parallel. The cartridge may further include one or more removable sample collection chambers suitable for interface with a standalone PCR thermal cycler or sequencing instrument. The cartridge itself may be suitable for interface with a standalone PCR thermal cycler or sequencing instrument. As used in this disclosure, the term “cartridge” may mean any assembly of parts containing samples and beads at the time of assay execution.

[0297] The cartridge may further include components designed to form physical or chemical barriers to prevent the diffusion of large molecules in order to minimize cross-contamination between microwells (or to increase path length and diffusion time). Examples of such barriers include, but are not limited to, a pattern of serpentine channels used for the delivery of cells and solid carriers (e.g., beads) to the microwell array; a retractable platen or deformable membrane that is pressed during a dissolution or incubation step to come into contact with the surface of the microwell array substrate; the use of larger beads, such as the Sephadex beads described above, to block the opening of microwells; or the release of an immiscible hydrophobic fluid from a reservoir in the cartridge during a dissolution or incubation step to effectively isolate and compartmentalize each microwell in the array.

[0298] In the cartridge design, the dimensions of the fluid channel and array chamber can be optimized to (i) provide uniform delivery of cells and beads to the microwell array, and (ii) minimize sample and reagent consumption. The width of the fluid channel can be 50 micrometers to 20 mm. In other embodiments, the width of the fluid channel can be at least 50 micrometers, at least 100 micrometers, at least 200 micrometers, at least 300 micrometers, at least 400 micrometers, at least 500 micrometers, at least 750 micrometers, at least 1 mm, at least 2.5 mm, at least 5 mm, at least 10 mm, or at least 20 mm. In yet another embodiment, the width of the fluid channel can be at most 20 mm, at most 10 mm, at most 5 mm, at most 2.5 mm, at most 1 mm, at most 750 micrometers, at most 500 micrometers, at most 400 micrometers, at most 300 micrometers, at most 200 micrometers, at most 100 micrometers, or at most 50 micrometers. The width of the fluid channel is approximately 2 mm. The width of the fluid channel can fall within any range bounded by any of these values ​​(for example, approximately 250 μm to approximately 3 mm).

[0299] The fluid channels in the cartridge may have depth. In the cartridge design, the depth of the fluid channels may be between 50 micrometers and 2 mm. The depth of the fluid channels may be at least 50 micrometers, at least 100 micrometers, at least 200 micrometers, at least 300 micrometers, at least 400 micrometers, at least 500 micrometers, at least 750 micrometers, at least 1 mm, at least 1.25 mm, at least 1.5 mm, at least 1.75 mm, or at least 2 mm. The depth of the fluid channels may be at most 2 mm, at most 1.75 mm, at most 1.5 mm, at most 1.25 mm, at most 1 mm, at most 750 micrometers, at most 500 micrometers, at most 400 micrometers, at most 300 micrometers, at most 200 micrometers, at most 100 micrometers, or at most 50 micrometers. The depth of the fluid channels is approximately 1 mm. The depth of the fluid channel can fall within any range bounded by any of these values ​​(for example, from about 800 micrometers to about 1 mm).

[0300] Cartridges can be manufactured using a variety of techniques and materials known to those skilled in the art. Generally, cartridges are manufactured as a series of individual components and then assembled using one of several mechanical assembly or joining techniques. Examples of preferred manufacturing techniques, but not limited to, include conventional machining, CNC machining, injection molding, thermoforming, and 3D printing. After the cartridge components are manufactured, they can be mechanically assembled using screws, clips, etc., or permanently bonded using one of various techniques (depending on the choice of materials used), for example, by heat bonding / welding, or by one of various adhesives or adhesive films, for example, epoxy, acrylic, silicone, UV-curable, polyurethane, or cyanoacrylate adhesives.

[0301] Cartridge components can be fabricated using any of several suitable materials, for example, but not limited to, silicon, fused silica, glass, any of various polymers, such as polydimethylsiloxane (PDMS, elastomer), polymethyl methacrylate (PMMA), polycarbonate (PC), polypropylene (PP), polyethylene (PE), high-density polyethylene (HDPE), polyimide, cyclic olefin polymer (COP), cyclic olefin copolymer (COC), polyethylene terephthalate (PET), epoxy resin, non-adherent materials such as Teflon® (PTFE), metals (e.g., aluminum, stainless steel, copper, nickel, chromium, and titanium), or any combination thereof.

[0302] The cartridge inlet and outlet features can be designed to provide a convenient and leak-free fluid connection with the instrument, or they can function as open reservoirs for manual pipetting of samples and reagents into or from the cartridge. Examples of convenient mechanical designs for the inlet and outlet port connectors include, but are not limited to, threaded connectors, Luer lock connectors, Luer slip or "slip tip" connectors, and press-fit connectors. The cartridge inlet and outlet ports may further include caps, spring-loaded covers or closures, or polymer membranes that may create an opening or perforation when the cartridge is placed in the instrument, and serve to prevent contamination of the internal cartridge surface during storage, or prevent fluid leakage when the cartridge is removed from the instrument. One or more outlet ports of the cartridge may further include removable sample collection chambers suitable for interface with a standalone PCR thermal cycler or sequencing instrument.

[0303] The cartridge may include an integrated miniature pump or other fluid-driven mechanism for controlling the fluid flow through the device. Suitable miniature pumps or fluid-driven mechanisms include, but are not limited to, electromechanical or pneumatically operated miniature syringe or plunger mechanisms, pneumatically operated or externally piston-operated membrane diaphragm pumps, pneumatically operated reagent pouches or bladders, or electroosmotic pumps.

[0304] The cartridge may include miniature valves for compartmentalizing preloaded reagents or controlling fluid flow through the device. Suitable examples of miniature valves include, but are not limited to, one-shot "valves" fabricated with molten or dissolvable wax or polymer plugs or perforated polymer membranes; pinch valves fabricated with deformable membranes and pneumatic, magnetic, electromagnetic, or electromechanical (solenoid) actuation; one-way valves fabricated with deformable membrane flaps; and miniature gate valves.

[0305] The cartridge may include a vent that provides an escape route for trapped air. The vent can be fabricated by various techniques, for example, using a porous plug made of polydimethylsiloxane (PDMS) or other hydrophobic material that allows for capillary wicking of air but blocks water penetration.

[0306] The cartridge's mechanical interface features are easily removable but can provide highly accurate and repeatable placement of the cartridge relative to the instrument system. Suitable mechanical interface features include, but are not limited to, alignment pins, alignment guides, and mechanical stops. Mechanical design features may include relief features that bring external devices, such as magnets or optical components, very close to the microwell array chamber.

[0307] The cartridge may also include a temperature control component or thermal interface feature paired with an external temperature control module. Suitable temperature control elements include, but are not limited to, resistance heating elements, miniature infrared light sources, Peltier heating or cooling devices, heat sinks, thermistors, and thermocouples. The thermal interface feature can be fabricated from a material with good thermal conductivity (e.g., copper, gold, silver, etc.) and may include one or more flat surfaces that allow for good thermal contact with an external heating or cooling block.

[0308] The cartridge may include an optical interface feature for use in optical imaging or spectroscopic examination of the microwell array. The cartridge may include an optically transparent window, for example, on the microwell substrate itself or on the side of the flow cell or microarray chamber facing the microwell array, and is made from a material that satisfies the spectral requirements of the imaging or spectroscopic technique used to probe the microwell array. Examples of suitable optical window materials, but not limited to, include glass, fused silica, polymethyl methacrylate (PMMA), polycarbonate (PC), cyclic olefin polymer (COP), or cyclic olefin copolymer (COC).

[0309] In at least some of the embodiments described above, one or more elements used in the embodiments are interchangeable in other embodiments, provided that such interchangeability is technically feasible. Those skilled in the art will see that various other omissions, additions, and modifications can be made to the methods and structures described above without departing from the scope of the claimed subject matter. All such modifications and changes are intended to fall within the scope of the subject matter defined in the appended claims.

[0310] In connection with the use of substantially any plural and / or singular terms as set forth herein, a person skilled in the art can convert from plural to singular and / or singular to plural, where appropriate in context and / or application. For clarity, various singular / plural substitutions may be explicitly described herein. Where used herein and in the appended claims, unless otherwise explicitly stated in context, the singular “a,” “an,” and “the” encompass multiple reference words. Any meaning of “or” herein is intended to encompass “and / or” unless otherwise specified.

[0311] Generally, it will be understood by those skilled in the art that the terminology used herein, particularly in the appended claims (for example, in the text of the appended claims), is generally intended to be "open" terminology (for example, the term "including" should be interpreted as "including but not limited to," the term "having" should be interpreted as "having at least," and the term "includes" should be interpreted as "including but not limited to."). Furthermore, if a specific number of introductory claim recitations is intended, such intention will be explicitly recited in the claims, and it will be understood by those skilled in the art that such intention does not exist in the absence of such recitations. For example, to aid in understanding, the following appended claims may include the use of the introductory phrases "at least one" and "one or more" to introduce claim recitations. However, even if such phrases are used, the introduction of a claim recitative with the indefinite article "a" or "an" should not be interpreted as meaning that any particular claim containing such an introducing claim recitative is limited to only one embodiment containing such recitative. This should not be interpreted even if the same claim contains the introducing phrase "one or more" or "at least one" and the indefinite article, for example, "a" or "an" (for example, "a" and / or "an" should be interpreted as meaning "at least one" or "one or more"). The same applies w...

Claims

1. A method for affixing barcodes to target polynucleotides in multiple partitions using a set of combinatorial barcodes, wherein the method is: A step of introducing several component barcodes to each of the aforementioned multiple partitions, Each component barcode includes a barcode subunit array and one or two linker arrays or their complements. The combinatorial barcode includes a first component barcode containing a cell label, selected from a first set of barcode subunit sequences, and a second component barcode containing a molecular label, selected from a second set of barcode subunit sequences. At least one of the component barcodes is fixed on a solid carrier, and each of the component barcodes fixed on the solid carrier includes a cell label. The process is as follows: Each partition in the plurality of partitions contains one cell and one solid carrier, and each partition is a droplet or microwell. and A step of forming a combinatorial barcode by connecting two or more component barcodes in each of the aforementioned plurality of partitions in the presence of a target polynucleotide of a sample, wherein the step is carried out by extending and / or amplifying the two or more component barcodes and the target polynucleotide to generate a transcript and / or amplicon containing the sequences of the two or more component barcodes, The process is as follows: The two or more component barcodes include a first component barcode and a second component barcode, the first component barcode includes a first barcode subunit sequence, the second component barcode includes a second barcode subunit sequence, the first barcode subunit sequence is a cell label, and the second barcode subunit sequence is a molecular label, thereby the combinatorial barcode includes a probability barcode containing a molecular label and a cell label, probability barcodes immobilized on the same solid carrier include the same cell label, and probability barcodes immobilized on different solid carriers include different cell labels. Methods that include...

2. The method according to claim 1, wherein the target polynucleotide comprises DNA and / or RNA.

3. The method according to any one of claims 1 to 2, wherein the combinatorial barcode includes two or more barcode subunit sequences linked to each other through a linker sequence and / or a target nucleotide sequence.

4. The method according to claim 3, wherein the two or more barcode subunit arrays form the combinatorial barcode by connecting to one or more linker arrays through hybridization and subsequent extension of the linker arrays.

5. The combinatorial barcode has an array of m barcode subunits, The method according to any one of claims 1 to 4, wherein, if necessary, the combinatorial barcode has m-1 or m-2 linker sequences, where m is an integer of 2 or more.

6. The aforementioned combinatorial barcode, The method according to claim 5, comprising an oligonucleotide containing the formula: barcode subunit sequence a-linker 1-barcode subunit sequence b-linker 2-... barcode subunit sequence c-linker m-1-barcode subunit sequence d.

7. The aforementioned combinatorial barcode, Formula: A first oligonucleotide comprising barcode subunit sequence a-linker 1-barcode subunit sequence b-linker 2-...-barcode subunit sequence c, and a second oligonucleotide comprising barcode subunit sequence d-...-linker 3-barcode subunit sequence e-linker m-2-barcode subunit sequence f, Includes, The method according to claim 6, wherein the first oligonucleotide and the second oligonucleotide may be linked to each other by a target nucleotide sequence, if necessary.

8. At least one of the two or more component barcodes includes a target-specific region. The method according to any one of claims 1 to 7, wherein the target-specific region optionally includes an oligo dT, a gene-specific sequence, or a random multimer sequence.

9. The method according to any one of claims 1 to 8, wherein the set of combinatorial barcodes includes at least 1,000, at least 10,000, at least 100,000, at least 200,000, at least 300,000, at least 400,000, at least 500,000, at least 1,000,000, at least 10,000,000, at least 100,000,000, and at least 1,000,000,000 or more unique combinatorial barcodes.

10. The method according to any one of claims 1 to 9, further comprising the step of hybridizing one of the two or more component barcodes in each partition to the target polynucleotide in the partition.

11. The method according to claim 10, wherein the component barcode hybridized to the target polynucleotide is used as a primer for the extension reaction to produce an extension product comprising the target and the combinatorial barcode.

12. The method according to claim 11, further comprising the step of amplifying and sequencing the extension product.

13. The process further includes pooling the contents of the plurality of partitions after extension and / or amplification, The method according to any one of claims 10 to 12, wherein the pool contents are sequenced as necessary.

14. The method according to any one of claims 1 to 13, wherein the solid carrier is a bead.

15. The method according to any one of claims 1 to 14, wherein the plurality of partitions are a plurality of microwells in a microwell array, and each microwell in the microwell array contains one solid carrier.