High-throughput single nucleus and single cell libraries and methods of making and using

JP2026035794A5Pending Publication Date: 2026-07-30ILLUMINA INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ILLUMINA INC
Filing Date
2025-12-03
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Traditional high-throughput screening (HTS) methods overlook subtle cellular changes and gene expression due to high costs and cellular heterogeneity, limiting mechanistic insights and missing rare subpopulations that survive chemotherapy.

Method used

Single-nucleus sequencing methods using hashing and normalization oligos to generate indexed libraries, enhancing sample throughput and reducing technical noise through doublet detection and sensitivity improvements.

Benefits of technology

Enables high-throughput analysis of cellular states and gene expression with reduced costs, overcoming cellular heterogeneity and improving sensitivity and specificity in sequencing data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000105_0000
    Figure 00000105_0000
  • Figure 00000105_0001
    Figure 00000105_0001
  • Figure 00000105_0002
    Figure 00000105_0002
Patent Text Reader

Abstract

High-throughput single nucleus and single cell libraries, and methods for production and use, are provided. [Solution] Provided herein are methods for preparing a sequencing library containing nucleic acids from multiple single cells. In one embodiment, the method includes nuclear or cell hashing, which allows for increased sample throughput at high collision rates and increased doublet detection. In one embodiment, the method includes normalized hashing, which helps estimate and remove technical noise in cell-to-cell variability, increasing sensitivity and specificity.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Field

[0002] Embodiments of the present disclosure relate to nucleic acid sequencing. Specifically, embodiments of the methods and compositions provided herein relate to using hashing oligos and / or normalization oligos to generate indexed single-nucleus and single-cell libraries and obtain sequence data therefrom.

[0003] (CROSS-REFERENCE TO RELATED APPLICATIONS)

[0004] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 812,853, filed March 1, 2019, which is incorporated herein by reference in its entirety.

[0005] (Sequence Listing)

[0006] This application contains a Sequence Listing that has been submitted electronically to the U.S. Patent and Trademark Office via EFS-Web as an ASCII text file “IP-1815-PCT_ST25.txt” 4 kilobytes in size, created on February 28, 2020. The information contained in the Sequence Listing is incorporated herein by reference.

[0007] (Government investment)

[0008] This invention was made with government support under Grant Nos. HG007811, HD088158, and R01 HG006283 awarded by the National Institutes of Health, and Grant No. DGE1258485 awarded by the National Science Foundation. The government has certain rights in this invention.

[0009] [Background technology]

[0010] High-throughput screening (HTS) is a cornerstone of the pharmaceutical drug discovery pipeline (J.R. Broach, J. Thorner, Nature, Vol. 384 (Suppl.), pp. 14-16 (1996); Pereira, J.A. Williams, Br. J. Pharmacol, 152, 53-61 (2007)). However, traditional HTS has at least two major limitations. First, most reads are limited to gross cellular phenotypes such as proliferation (D. Shum et al., J. Enzyme Inhib. Med. Chem. 23, 931-945 (2008); C. Yu et al., Nat. Biotechnol. 34, 419-423 (2016)), morphology (Z. E. Perlman et al., Science 306, 1194-1198 (2004); Y. Futamura et al., Chem. Biol. 19, 1620-1630 (2012)), or highly specific molecular reads (J. Kang et al., Nat. Biotechnol. 34, 70-77 (2016); K. L. Huss, P. E. Blonigen, R. M. Campbell, J. Biomol. Screen. 12, 578-584 (2007)). Subtle changes in cellular state and gene expression that could otherwise provide mechanistic insights or reveal off-target effects are routinely overlooked.

[0011] Second, even when HTS is performed in conjunction with more comprehensive molecular phenotyping such as transcriptional profiling (C. Ye et al., Nat. Commun. 9, 4307 (2018); E. C. Bush et al., Nat. Commun. 8, 105 (2017); A. Subramanian et al., Cell 171, 1437-1452.e17 (2017); Lamb et al., Science 313, 1929-1935 (2006)), a limitation of bulk assays is that even cells of ostensibly the same "type" may exhibit heterogeneous responses (M. B. Elowitz, A. J. Levine, E. D. Siggia, P. S. Wain, Science 297, 1183-1186 (2002); C. Trapnell, Genome Res. 25, 1491-1498 (2015)). Such cellular heterogeneity is highly relevant in vivo. For example, it is largely unknown whether rare subpopulations of cells that survive chemotherapy do so based on their genetic background, epigenetic state, or some other aspect (S.M. Shaffer et al., Nature 546, 431-435 (2017); S.J. Spencer, S. Gaudet, J.G. Albeck, J.M. Burke, P.K. Sorger, Nature 459, 428-432 (2009)). Furthermore, the level of sparsity and technical noise often makes it difficult to extract biologically meaningful information.

[0012] In principle, single-cell transcriptome sequencing (scRNA-seq) represents a form of high-content molecular phenotyping that may enable HTS to overcome both limitations. However, the per-sample and per-cell costs of most scRNA-seq techniques remain high, precluding even moderately sized screens. Recently, "cell hashing" methods have been developed, in which cells from different samples are molecularly labeled and mixed prior to scRNA-seq. However, current hashing methods require relatively expensive reagents (e.g., antibodies (M. Stoeckius et al., Genome Biol. 19, 224 (2018)) or chemically modified DNA oligos (J. Gehring, J.H. Park, S. Chen, M. Thomson, L. Pachter, bioRxiv 315333 [preprint] 5 May 2018, doi.org / 10.1101 / 315333, C.S. McGinnis et al., Nat. Methods 16, 619-626 (2019))), use cell-type-dependent protocols (D. Shin, W. Lee, J.H. Lee, D. Bang, Sci. Adv. 5, eaav2249 (2019)), and / or use scRNA-seq platforms that have a high cost per cell. [Prior art documents] [Non-patent literature]

[0013] [Non-Patent Document 1] J.R. Broach, J. Thorner, Nature, Vol. 384 (Suppl.), 14-16 (1996) [Non-patent document 2] Pereira, JA Williams, Br.J.Pharmacol, 152,53-61(2007) [Non-patent document 3] D. Shum et al., J. Enzyme Inhib. Med. Chem. 23, 931-945 (2008) [Non-patent document 4] C.Yuら, Nat.Biotechnol.34, 419-423 (2016) ZEPerlmanら, Science 306, 1194-1198 (2004)

Non-licensed Document 5

Non-licensed Document 6

Non-licensed Document 7

Non-licensed Document 8

Non-licensed literature 9

Non-licensed literature 10

Non-licensed Document 11

Non-licensed Document 12

Non-licensed Document 13

Non-licensed Document 14

Non-licensed Document 15

Non-licensed Document 16

[0014] Single-cell and single-nucleus sequencing of high cell numbers using single-nucleus sequencing (sci-) methods has demonstrated its effectiveness in separating populations within cells and complex tissues through differences in transcriptome, chromatin accessibility, mutational differences, and other differences. One method described herein, nuclear hashing or cell hashing, uses hashing oligos to increase sample throughput and doublet detection at high collision rates. Another method described herein, normalized hashing, uses normalization oligos as a standard to help estimate and remove technical noise in cell-to-cell variation, increasing sensitivity and specificity.

[0015] A method for preparing a sequencing library is provided herein. In one embodiment, the library contains nucleic acids from a plurality of single nuclei or single cells, and the method includes providing the plurality of cells in a first plurality of compartments and contacting isolated nuclei from the cells in each compartment or the cells in each compartment with a hashing oligo to generate hashed nuclei or hashed cells. In one embodiment, at least one copy of the hashing oligo is associated with the isolated nuclei or cells. In one embodiment, the hashing oligo comprises a hashing index. In one embodiment, the hashing index in each compartment comprises an index sequence that is different from the index sequences in other compartments. The association between the hashing oligo and the isolated nuclei or cells can be non-specific, such as by absorption. The method can further include combining hashed nuclei or hashed cells from different compartments to generate pooled hashed nuclei or pooled hashed cells. In one embodiment, the method can further include exposing the plurality of cells in each compartment to a predetermined condition. The exposure to the predetermined condition can be at any time during the method and, in one embodiment, occurs before the contacting.

[0016] In one embodiment, the method can optionally include processing the pooled hashed cells or pooled hashed nuclei using a single-cell combinatorial indexing method to obtain a sequencing library containing nucleic acids from a plurality of single nuclei. Examples of single-cell combinatorial indexing methods that can be used can include, but are not limited to, single-nucleus transcriptome sequencing, single-cell transcriptome sequencing, single-cell transcriptome and transposon-accessible chromatin sequencing, single-nucleus whole genome sequencing, single-nucleus sequencing of transposon-accessible chromatin, sci-HiC, drug-seq, sci-CAR, sci-MET, sci-Crop, sci-perturb, or sci-Crispr.

[0017] A method for normalizing a sequencing library is also provided by the present disclosure. In one embodiment, the sequencing library contains nucleic acids from a plurality of single nuclei or single cells. In one embodiment, the method includes providing a first plurality of compartments containing isolated nuclei or cells and contacting the isolated nuclei or cells in each compartment with a population of normalizing oligos, wherein members of each population of normalizing oligos are associated with the isolated nuclei or cells. In one embodiment, the contacting occurs before the isolated nuclei or cells are distributed into compartments. The normalizing oligos can be associated with the isolated nuclei or cells before or after compartmentalization. The association between the normalizing oligos and the isolated nuclei or cells can be non-specific, such as by absorption. The method may further include combining labeled nuclei or labeled cells from different compartments to generate pooled labeled nuclei or pooled labeled cells. In one embodiment, the method may further include exposing a plurality of cells in each compartment to a predetermined condition. The exposure to the predetermined condition can be at any time during the method and, in one embodiment, occurs before the contacting.

[0018] definition

[0019] Terms used herein will be understood to take their ordinary meaning in the relevant art unless otherwise specified. Some terms used herein and their meanings are set forth below.

[0020] As used herein, the terms "organism" and "subject" are used interchangeably and refer to microorganisms (e.g., prokaryotes or eukaryotes), animals, and plants. An example of an animal is a mammal, such as a human.

[0021] As used herein, the term "cell type" is intended to identify cells based on morphology, phenotype, developmental origin, or other known or recognizable distinguishing cellular characteristics. A variety of different cell types can be obtained from a single organism (or from organisms of the same species). Exemplary cell types include gametes (e.g., including female gametes, e.g., eggs or ovum cells, and male gametes, e.g., sperm), ovarian epithelium, ovarian epithelial cells, ovarian fibroblasts, testis, bladder, pancreatic epithelium, pancreatic alpha cells, immune cells, B cells, T cells, natural killer cells, dendritic cells, cancer cells, eukaryotic cells, stem cells, blood cells, muscle cells, adipocytes, skin cells, nerve cells, bone cells, pancreatic cells, endothelial cells, pancreatic beta cells, pancreatic endothelium, bone marrow lymphoblasts, bone marrow lymphoblasts, bone marrow macrophages, myeloblasts, bone marrow adipocytes, bone marrow osteoblasts, bone marrow chondrocytes, bone marrow chondrocytes, myeloblasts, bone marrow chondrocytes, promyeloblasts, bone marrow megakaryoblasts, bladder, brain B lymphocytes, brain glial cells, neurons, brain astrocytes, neuroectoderm, brain macrophages, brain microglia, brain epithelium, cortical neurons, brain fibroblasts, breast epithelium, colon epithelium, Intestinal B lymphocytes, mammary epithelium, mammary myoepithelium, mammary fibroblasts, colon enterocytes, cervical epithelium, mammary ductal epithelium, tongue epithelium, tonsillar dendritic cells, tonsillar lymphocytes, peripheral blood lymphoblasts, peripheral blood T lymphoblasts, peripheral blood T lymphocytes, peripheral blood natural killer cells, peripheral blood B lymphoblasts, peripheral blood monocytes, peripheral blood myeloblasts, peripheral blood monoblasts, peripheral blood monoblasts, peripheral blood monoblasts, peripheral blood monoblasts, peripheral blood T lymphocytes , peripheral blood promyeloblasts, peripheral blood macrophages, peripheral blood basophils, liver endothelium, liver mast, liver epithelium, liver B lymphocytes, splenic endothelium, splenic epithelium, splenic B lymphocytes, hepatocytes, liver, fibroblasts, lung epithelium, bronchial epithelium, lung fibroblasts, lung B lymphocytes, lung Schwann, lung squamous epithelium, lung macrophages, lung osteoblasts, neuroendocrine, lung alveolar, gastric epithelium, and gastric fibroblasts.

[0022] As used herein, the term "tissue" is intended to mean a collection or aggregation of cells that act together to perform one or more specific functions within an organism. Cells may optionally be morphologically similar. Exemplary tissues include, but are not limited to, embryonic, epididymal, eye, muscle, skin, tendon, vein, artery, blood, heart, spleen, lymph node, bone, bone marrow, lung, bronchus, trachea, intestine, small intestine, large intestine, colon, rectum, salivary gland, tongue, gallbladder, appendix, liver, pancreas, brain, stomach, skin, kidney, ureter, bladder, urethra, gonads, testicle, ovary, uterus, fallopian tube, thymus, pituitary gland, thyroid gland, adrenal gland, or parathyroid gland. Tissues may be derived from any of a variety of organs in humans or other organisms. Tissues may be healthy or unhealthy. Examples of unhealthy tissues include malignant tumors of reproductive tissue, lung, breast, colorectal, prostate, nasopharynx, stomach, testes, skin, nervous system, bone, ovary, liver, blood tissue, pancreas, uterus, kidney, lymphatic tissue, etc. Malignant tumors can be of various histological subtypes, such as, but not limited to, carcinoma, adenocarcinoma, sarcoma, fibroadenocarcinoma, neuroendocrine, or undifferentiated.

[0023] As used herein, the term "compartment" is intended to mean an area or volume that separates or isolates something from another. Exemplary compartments include, but are not limited to, vials, tubes, wells, droplets, boluses, beads, containers, surface features, or areas or volumes separated by physical forces such as fluid flow, magnetism, or electric current. In one embodiment, a compartment is a well of a multiwell plate, such as a 96- or 384-well plate. As used herein, a droplet may include hydrogel beads, which are beads for encapsulating one or more nuclei or cells and include a hydrogel composition. In some embodiments, the droplet is a homogenous droplet of hydrogel material or a hollow droplet with a polymer hydrogel shell. Whether homogenous or hollow, the droplet may be capable of encapsulating one or more nuclei or cells.

[0024] As used herein, a "transposome complex" refers to a nucleic acid comprising an integrase and an integration recognition site. A "transposome complex" is a functional complex formed by a transposase and a transposase recognition site capable of catalyzing a transposition reaction (see, for example, Gunderson et al., WO 2016 / 130704). Examples of integrases include, but are not limited to, integrases or transposases. Examples of integration recognition sites include, but are not limited to, transposase recognition sites.

[0025] As used herein, the term "nucleic acid" is intended to be consistent with its use in the art and includes naturally occurring nucleic acids or functional analogs thereof. Particularly useful functional analogs are capable of hybridizing to nucleic acids in a sequence-specific manner or can be used as templates for replicating specific nucleotide sequences. Naturally occurring nucleic acids generally have backbones containing phosphodiester bonds. Analog structures can have alternative backbone linkages, including any of a variety known in the art. Naturally occurring nucleic acids generally have deoxyribose sugars (e.g., found in deoxyribonucleic acid (DNA)) or ribose sugars (e.g., found in ribonucleic acid (RNA)). Nucleic acids can contain any of a variety of analogs of these sugar moieties known in the art. Nucleic acids can include natural or unnatural bases. In this regard, natural deoxyribonucleic acids can have one or more bases selected from the group consisting of adenine, thymine, cytosine, or guanine, and ribonucleic acids can have one or more bases selected from the group consisting of adenine, uracil, cytosine, or guanine. Useful unnatural bases that can be included in nucleic acids are known in the art. Examples of unnatural bases include locked nucleic acids (LNAs), bridged nucleic acids (BNAs), and pseudo-complementary bases (Trilink Biotechnologies, San Diego, California). LNA and BNA bases can be incorporated into DNA oligonucleotides to increase the hybridization strength and specificity of the oligonucleotides. LNA and BNA bases, and the use of such bases, are known and routine to those skilled in the art.

[0026] As used herein, the term "target," when used in reference to a nucleic acid, is intended as a semantic identifier of the nucleic acid in the context of a method or composition described herein and does not necessarily limit the structure or function of the nucleic acid beyond what is otherwise explicitly indicated. A target nucleic acid can be essentially any nucleic acid of known or unknown sequence. It can be, for example, a fragment of genomic DNA (e.g., chromosomal DNA), a plasmid, cell-free DNA, RNA (e.g., mRNA), a protein (e.g., a cellular or cell surface protein), or extrachromosomal DNA such as cDNA. Sequencing can result in determining the sequence of all or part of a target molecule. Targets can be derived from primary nucleic acid samples, such as nuclei. In one embodiment, targets can be processed into templates suitable for amplification by placing universal sequences at the ends of each target fragment. Targets can also be obtained from primary RNA samples by reverse transcription into cDNA. In one embodiment, target is used in reference to a subset of DNA, RNA, or proteins present in a cell. Target sequencing typically uses the selection and isolation of a gene, region, or protein of interest, either by PCR amplification (e.g., region-specific primers) or hybridization-based capture methods or antibodies. Target enrichment can be performed at various stages of the method. For example, target RNA representation can be achieved by using target-specific primers in the reverse transcription step or by hybridization-based enrichment of subsets from more complex libraries. Examples include exome sequencing or the L1000 assay (Subramanian et al., 2017, Cell, 171;1437-1452). Target sequencing can involve any enrichment process known to those skilled in the art.

[0027] As used herein, the term "universal," when used to describe a nucleotide sequence, refers to a region of sequence common to two or more nucleic acid molecules or samples, where the molecules also have regions of sequence that differ from one another. Universal sequences present in different members of a population of molecules can be used to capture multiple different nucleic acids using a population of universal capture nucleic acids, e.g., capture oligonucleotides complementary to a portion of the universal sequence, e.g., universal capture sequences. Non-limiting examples of universal capture sequences include sequences identical to or complementary to P5 and P7 primers. Similarly, universal sequences present in different members of a population of molecules can be used to replicate (e.g., sequence) or amplify multiple different nucleic acids using a population of universal primers complementary to a portion of the universal sequence, e.g., universal anchor sequences. In one embodiment, the universal anchor sequence is used as the site to which the universal primer (e.g., sequencing primer for read 1 or read 2) anneals for sequencing. Thus, the capture oligonucleotide or universal primer comprises a sequence that can specifically hybridize to the universal sequence.

[0028] The terms "P5" and "P7" may be used to refer to universal capture sequences or capture oligonucleotides. The terms "P5" (P5 prime) and "P7" (P7 prime) refer to the complements of P5 and P7, respectively. It will be understood that any suitable universal capture sequence or capture nucleotide can be used in the methods presented herein, and the use of P5 and P7 is only an exemplary embodiment. The use of capture nucleotides such as P5 and P7 or their complements on flow cells is known in the art, as exemplified by the disclosures of WO 2007 / 010251, WO 2006 / 064199, WO 2005 / 065814, WO 2015 / 106941, WO 1998 / 044151, and WO 2000 / 018957. For example, any suitable forward amplification primer, whether immobilized or in solution, can be useful in the methods provided herein for amplifying complementary sequences and sequences. Similarly, any suitable reverse amplification primer, whether immobilized or in solution, can be useful in the methods provided herein for amplifying complementary sequences and sequences. Those skilled in the art will understand how to design and use suitable primer sequences for capturing and / or amplifying nucleic acids as provided herein.

[0029] As used herein, the term "primer" and its derivatives generally refer to any nucleic acid capable of hybridizing to a target sequence of interest. Typically, a primer functions as a substrate onto which nucleotides can be polymerized by a polymerase or to which nucleotides can be ligated; however, in some embodiments, a primer can be incorporated into a synthesized nucleic acid strand to provide a site to which another primer can hybridize and prime synthesis of a new strand complementary to the synthesized nucleic acid molecule. A primer can contain any combination of nucleotides or their analogs. In some embodiments, a primer is a single-stranded oligonucleotide or polynucleotide. The terms "polynucleotide" and "oligonucleotide" are used interchangeably herein to refer to polymeric forms of nucleotides of any length and can include ribonucleotides, deoxyribonucleotides, their analogs, or mixtures thereof. It should be understood that these terms include, as equivalents, analogs of any of DNA, RNA, cDNA, or antibody-oligoconjugates made from nucleotide analogs, and are applicable to single-stranded (such as sense or antisense) and double-stranded polynucleotides. As used herein, the term also encompasses cDNA, which is complementary or copy DNA produced from an RNA template, for example, by the action of reverse transcriptase. The term refers only to the primary structure of the molecule. Thus, the term includes triple-, double-, and single-stranded deoxyribonucleic acid ("DNA"), as well as triple-, double-, and single-stranded ribonucleic acid ("RNA").

[0030] As used herein, the term "adapter" and its derivatives, such as universal adapter, generally refer to any linear oligonucleotide that can be attached to a nucleic acid molecule of the present disclosure. In some embodiments, the adapter is substantially non-complementary to the 3' or 5' end of any target sequence present in a sample. In some embodiments, suitable adapter lengths range from about 10-100 nucleotides, about 12-60 nucleotides, or about 15-50 nucleotides in length. Generally, an adapter can comprise any combination of nucleotides and / or nucleic acids. In some aspects, an adapter can comprise one or more cleavable groups at one or more positions. In other aspects, an adapter can comprise a sequence that is substantially identical to or substantially complementary to at least a portion of a primer, such as a universal primer. In some embodiments, an adapter can comprise a barcode (also referred to herein as a tag or index) to assist in downstream error correction, identification, or sequencing. The terms "adaptor" and "adapter" are used interchangeably.

[0031] As used herein, the term "each," when used in reference to a collection of items, is intended to identify each individual item in the set, but does not necessarily refer to every item in the set, unless the context clearly dictates otherwise.

[0032] As used herein, the term "transport" refers to the movement of molecules through a fluid. The term can include passive transport, such as the movement of molecules along their concentration gradient (e.g., passive diffusion). The term can also include active transport, in which molecules can move along or against their concentration gradient. Thus, transport can include the application of energy to move one or more molecules in a desired direction or to a desired location, such as an amplification site.

[0033] As used herein, "amplification," "amplifying," or "amplification reaction," and derivatives thereof, generally refer to any act or process in which at least a portion of a nucleic acid molecule is duplicated or copied onto at least one additional nucleic acid molecule. The additional nucleic acid molecule optionally comprises a sequence that is substantially identical to or substantially complementary to at least a portion of a template nucleic acid molecule. The template nucleic acid molecule may be single-stranded or double-stranded, and the additional nucleic acid molecules may independently be single-stranded or double-stranded. Amplification optionally involves linear or exponential replication of nucleic acid molecules. In some embodiments, such amplification can be performed using isothermal conditions; in other embodiments, such amplification can involve thermal cycling. In some embodiments, amplification is a multiplex amplification that involves simultaneous amplification of multiple target sequences in a single amplification reaction. In some embodiments, "amplification" includes amplifying at least a portion of DNA- and RNA-based nucleic acids, alone or in combination. The amplification reaction can involve any amplification process known to those of skill in the art. In some embodiments, the amplification reaction involves polymerase chain reaction (PCR).

[0034] As used herein, "amplification conditions" and its derivatives generally refer to conditions suitable for amplifying one or more nucleic acid sequences. Such amplification can be linear or exponential. In some embodiments, amplification conditions can include isothermal conditions, or can include thermocycling conditions, or a combination of isothermal and thermocycling conditions. In some embodiments, conditions suitable for amplifying one or more nucleic acid sequences include polymerase chain reaction (PCR) conditions. Typically, amplification conditions refer to a reaction mixture sufficient to amplify a nucleic acid, such as one or more target sequences flanked by universal sequences, or amplified target sequences ligated to one or more adapters. Generally, amplification conditions include a catalyst for amplification, or nucleic acid synthesis, e.g., a polymerase, primers having a degree of complementarity to the nucleic acid to be amplified, and nucleotides, such as deoxyribonucleotide triphosphates (dNTPs), to facilitate primer extension when hybridized to the nucleic acid. Amplification conditions may require hybridization or annealing of a primer to a nucleic acid, extension of the primer, and a denaturation step in which the extended primer is separated from the nucleic acid sequence undergoing amplification. Typically, but not necessarily, amplification conditions may include thermal cycling, although in some embodiments, amplification conditions include multiple cycles in which the annealing, extension, and separation steps are repeated. Typically, amplification conditions include Mg 2+ or Mn 2+ and may also include various modifiers of ionic strength.

[0035] As used herein, "reamplification" and its derivatives generally refer to any process (in some embodiments referred to as a "secondary" amplification) in which at least a portion of an amplified nucleic acid molecule is further amplified via any suitable amplification process, thereby producing a re-amplified nucleic acid molecule. The secondary amplification need not be identical to the original amplification process by which the amplified nucleic acid molecule was produced, nor need the amplified nucleic acid molecule be completely identical to or completely complementary to the amplified nucleic acid molecule; all that is required is that the re-amplified nucleic acid molecule comprise at least a portion of the amplified nucleic acid molecule or its complement. For example, reamplification can involve different amplification conditions and / or the use of different primers, including different target-specific primers than the primary amplification.

[0036] As used herein, the term "polymerase chain reaction" ("PCR") refers to the Mullis method in U.S. Pat. Nos. 4,683,195 and 4,683,202, which describes a method for increasing the concentration of a polynucleotide segment of interest in a mixture of genomic DNA without cloning or purification. This process for amplifying a polynucleotide of interest involves introducing a large excess of two oligonucleotide primers into a DNA mixture containing the desired polynucleotide of interest, followed by a series of thermal cycling steps in the presence of a DNA polymerase. The two primers are complementary to each strand of the double-stranded polynucleotide of interest. The mixture is first denatured at a higher temperature, and then the primers are annealed to complementary sequences within the polynucleotide of the molecule of interest. After annealing, the primers are extended with a polymerase to form a new pair of complementary strands. The steps of denaturation, primer annealing, and polymerase extension can be repeated multiple times (called thermal cycling) to obtain a highly concentrated amplified segment of the desired polynucleotide of interest. The length of the amplified segment (amplicon) of the desired target polynucleotide is determined by the relative positions of the primers with respect to each other, and therefore this length is a controllable parameter. By repeating this process, the method is called PCR. Because the desired amplified segment of the target polynucleotide becomes the predominant nucleic acid sequence (in terms of concentration) in the mixture, it is said to be "PCR amplified." In a modification of the above method, the target nucleic acid molecule can be PCR amplified using multiple different primer pairs, and in some cases, one or more primer pairs can be used per target nucleic acid molecule of interest, thereby forming a multiplex PCR reaction.

[0037] As defined herein, "multiplex amplification" refers to the selective, non-random amplification of two or more target sequences in a sample using at least one target-specific primer. In some embodiments, multiplex amplification is performed such that some or all of the target sequences are amplified in a single reaction vessel. The "plex" of a given multiplex amplification generally refers to the number of different target-specific sequences amplified during that single multiplex amplification. In some embodiments, the plex can be about 12-plex, 24-plex, 48-plex, 96-plex, 192-plex, 384-plex, 768-plex, 1536-plex, 3072-plex, 6144-plex, or more. The amplified target sequences can be analyzed by several different methodologies (e.g., gel electrophoresis followed by densitometry, quantification by bioanalyzer or quantitative PCR, hybridization with labeled probes, incorporation of biotinylated primers followed by avidin-enzyme conjugate detection, or detection of the amplified target sequences). 32 Detection is also possible by incorporation of P-labeled deoxynucleotide triphosphates.

[0038] As used herein, "amplified target sequence" and its derivatives generally refer to a nucleic acid sequence produced by amplifying a target sequence using target-specific primers and the methods provided herein. The amplified target sequence may be either of the same sense (i.e., positive strand) or antisense (i.e., negative strand) with respect to the target sequence.

[0039] As used herein, the terms "ligate," "ligation," and derivatives thereof generally refer to the process of covalently linking two or more molecules to one another, e.g., covalently linking two or more nucleic acid molecules to one another. In some embodiments, ligation involves joining nicks between adjacent nucleotides of nucleic acids. In some embodiments, ligation involves forming a covalent bond between an end of a first nucleic acid molecule and an end of a second nucleic acid molecule. In some embodiments, ligation can involve forming a covalent bond between a 5' phosphate group of one nucleic acid and a 3' hydroxyl group of a second nucleic acid, thereby forming a ligated nucleic acid molecule. Generally, for purposes of this disclosure, an amplified target sequence can be ligated to an adapter to generate an adapter-ligated amplified target sequence.

[0040] As used herein, "ligase" and its derivatives generally refer to any agent capable of catalyzing the ligation of two substrate molecules. In some embodiments, a ligase includes an enzyme capable of catalyzing the joining of nicks between adjacent nucleotides of nucleic acids. In some embodiments, a ligase includes an enzyme capable of catalyzing the formation of a covalent bond between the 5' phosphate of one nucleic acid molecule and the 3' hydroxyl of another nucleic acid molecule, thereby forming a ligated nucleic acid molecule. Suitable ligases include, but are not limited to, T4 DNA ligase, T4 RNA ligase, and E. coli DNA ligase.

[0041] As used herein, "ligation conditions" and its derivatives generally refer to conditions suitable for ligating two molecules to one another. In some embodiments, the ligation conditions are suitable for sealing a nick or gap between nucleic acids. As used herein, the term nick or gap is consistent with the use of the term in the art. Typically, a nick or gap can be ligated in the presence of an enzyme, such as a ligase, at an appropriate temperature and pH. In some embodiments, T4 DNA ligase can bind to a nick between nucleic acids at a temperature of about 70-72°C.

[0042] As used herein, the term "flow cell" refers to a chamber containing a solid surface through which one or more fluid reagents can flow. Examples of flow cells and associated fluidic systems and detection platforms that can be readily used in the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497, U.S. Patent No. 7,057,026, WO 91 / 06678, WO 07 / 123744, U.S. Patent No. 7,329,492, U.S. Patent No. 7,211,414, U.S. Patent No. 7,315,019, U.S. Patent No. 7,405,281, and U.S. Patent No. 2008 / 0108082.

[0043] As used herein, the term "amplicon," when used with reference to a nucleic acid, refers to a product of copying a nucleic acid, which product has a nucleotide sequence identical to or complementary to at least a portion of the nucleotide sequence of the nucleic acid. Amplicons can be produced by any of a variety of amplification methods using a nucleic acid or its amplicon as a template, including, for example, polymerase extension, polymerase chain reaction (PCR), rolling circle amplification (RCA), ligation extension, or ligation chain reaction. An amplicon can be a nucleic acid molecule having a single copy of a particular nucleotide sequence (e.g., a PCR product) or multiple copies of a nucleotide sequence (e.g., a concatemeric product of RCA). A first amplicon of a target nucleic acid is typically a complementary copy. Subsequent amplicons are copies made from the target nucleic acid or the first amplicon after the generation of the first amplicon. Subsequent amplicons can have a sequence that is substantially complementary to or substantially identical to the target nucleic acid.

[0044] As used herein, the term "amplification site" refers to a site within or on a sequence where one or more amplicons can be generated. An amplification site can be further configured to contain, retain, or attach at least one amplicon generated at that site.

[0045] As used herein, the term "array" refers to a collection of sites that can be distinguished from one another according to their relative positions. Different molecules at different sites of an array can be distinguished from one another according to the site's position within the array. Each site of an array can contain one or more molecules of a particular type. For example, a site can contain a single target nucleic acid molecule having a particular sequence, or a site can contain several nucleic acid molecules having the same sequence (and / or its complementary sequence). The sites of an array can be different features located on the same substrate. Exemplary features include, but are not limited to, wells in a substrate, beads (or other particles) in or on a substrate, protrusions from a substrate, ridges on a substrate, or channels within a substrate. The sites of an array can be separate substrates, each with a different molecule. The different molecules attached to the separate substrates can be identified according to the position of the substrate on a surface to which the substrates are associated, or according to the position of the substrate within a liquid or gel. An exemplary array in which separate substrates are located on a surface includes, but is not limited to, beads in wells.

[0046] As used herein, the term "capacity," when used in reference to a site and nucleic acid material, refers to the maximum amount of nucleic acid material that can occupy the site. For example, the term can refer to the total number of nucleic acid molecules that can occupy the site under certain conditions. Other measurements can be used, including, for example, the total mass of nucleic acid material that can occupy the site under certain conditions or the total number of copies of a particular nucleotide sequence. Typically, the volume of a site for a target nucleic acid is substantially equivalent to the volume of the site for an amplicon of the target nucleic acid.

[0047] As used herein, the term "capture agent" refers to a material, chemical, molecule, or portion thereof that can attach to, retain, or bind to a target molecule (e.g., a target nucleic acid). Exemplary capture agents include, but are not limited to, a capture nucleic acid (also referred to herein as a capture oligonucleotide) that is complementary to at least a portion of a target nucleic acid, a member of a receptor-ligand binding pair (e.g., avidin, streptavidin, biotin, lectins, carbohydrates, nucleic acid-binding proteins, epitopes, antibodies, etc.) that can bind to a target nucleic acid (or a linking moiety attached thereto), or a chemical reagent that can form a covalent bond with a target nucleic acid (or a linking moiety attached thereto).

[0048] As used herein, the term "reporter moiety" can refer to any identifiable tag, label, index, barcode, or group that allows for determining the composition, identity, and / or source of an interrogated analyte. In some embodiments, the reporter moiety can include an antibody that specifically binds to a protein. In some embodiments, the antibody can include a detectable label. In some embodiments, the reporter can include an antibody or affinity reagent labeled with a nucleic acid tag. The nucleic acid tag can be detectable, for example, via proximity ligation assay (PLA) or proximity extension assay (PEA) or sequence-based reads (Shahi et al., Scientific Record Volume 7, Reference Number: 44447, 2017) or CITE-seq (Stoeckius et al., Nature Methods 14:865-868, 2017).

[0049] As used herein, the term "clonal population" refers to a population of nucleic acids that are homogeneous with respect to a particular nucleotide sequence. Homogeneous sequences are typically at least 10 nucleotides in length, but may be longer, e.g., at least 50, 100, 250, 500, or 1000 nucleotides in length. A clonal population can be derived from a single target or template nucleic acid. Typically, all nucleic acids in a clonal population have the same nucleotide sequence. It will be understood that minor variations (e.g., due to amplification artifacts) can occur without deviating from clonality.

[0050] As used herein, "providing" in the context of a composition, article, nucleic acid, or nucleus means making the composition, article, nucleic acid, or nucleus, purchasing the composition, article, nucleic acid, or nucleus, or otherwise obtaining the compound, composition, article, or nucleus.

[0051] As used herein, an "index" (also referred to as an "index region," "index adapter," "tag," or "barcode") refers to a unique nucleic acid tag that can be used to identify a sample or source of nucleic acid material. When nucleic acid samples are derived from multiple sources, the nucleic acids in each nucleic acid sample can be tagged with a different nucleic acid tag so that the source of the sample can be identified. Any suitable index or set of indexes can be used, as known in the art and as exemplified by the disclosures of U.S. Pat. No. 8,053,192, WO 05 / 068656, and U.S. Patent Application Publication No. 2013 / 0274117. In some embodiments, the index can include a 6-base Index1 (i7) sequence, an 8-base Index1 (i7) sequence, an 8-base Index2 (i5e) sequence, a 10-base Index1 (i7) sequence, or a 10-base Index2 (i5) sequence from Illumina (San Diego, CA).

[0052] As used herein, the term "unique molecular identifier" or "UMI" refers to a molecular tag that can be attached to a nucleic acid molecule, either randomly, non-randomly, or semi-randomly. When incorporated into a nucleic acid molecule, the unique molecular identifier (UMI) can be used to correct for subsequent amplification bias by directly counting the UMI, which is sequenced after amplification.

[0053] The term "and / or" means one or all of the listed elements or a combination of any two or more of the listed elements.

[0054] The words "preferred" and "preferably" refer to embodiments of the present disclosure that may offer certain benefits, under particular circumstances. However, other embodiments may also be preferred, under the same or other circumstances. Furthermore, the recitation of one or more preferred embodiments does not imply that other embodiments are not useful, and is not intended to exclude other embodiments from the scope of the present disclosure.

[0055] The term "comprises" and variations thereof mean that these terms are used in the description and claims. When appearing in the text, it does not have a limiting meaning.

[0056] As used herein, when words such as "include," "includes," or "including" are used herein, they also include similar embodiments described with the terms "consisting of" and / or "consisting essentially of." It is understood that embodiments are also provided.

[0057] Unless otherwise specified, "a," "an," "the," and "at least one" are used interchangeably and mean one or more than one.

[0058] As used herein, the recitations of numerical ranges by endpoints include all numbers subsumed within that range (eg, 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.80, 4, 5, etc.).

[0059] In any method disclosed herein that includes distinct steps, the steps may be performed in any practicable order, and, suitably, any combination of two or more steps may be performed simultaneously.

[0060] References to "one embodiment," "an embodiment," "particular embodiments," or "some embodiments" mean that the particular feature, configuration, composition, or characteristic described in connection with this embodiment is included in at least one embodiment of the present disclosure. Thus, the appearances of such phrases in various places throughout this specification do not necessarily refer to the same embodiment of the present disclosure. Furthermore, the particular features, configurations, compositions, or characteristics may be combined in any suitable manner in one or more embodiments.

[0061] [Brief explanation of the drawings]

[0062] The following detailed description of the exemplary embodiments of the present disclosure can be best understood when read in conjunction with the following drawings.

[0063] [Figure 1] FIG. 1 shows a general block diagram of a general exemplary method of one embodiment of nuclear or cell hashing according to the present disclosure.

[0064] [Figure 2] 1 shows a general block diagram of a general exemplary method of one embodiment of canonicalized hashing according to the present disclosure.

[0065] [Figure 3] 1 shows a general block diagram of a general exemplary method of one embodiment of a single-cell combinatorial indexing method by nuclear hashing according to the present disclosure.

[0066] [Figure 4A] Figure 1 shows that SqueePLEX uses polyadenylated single-stranded oligonucleotides to label nuclei, enabling cell hashing and doublet detection. Fluorescence images of permeabilized nuclei after incubation with DAP1 (top) and Alexa Fluor-647-conjugated single-stranded oligonucleotides (bottom). [Figure 4B] We demonstrate that sci-PLEX uses polyadenylated single-stranded oligonucleotides to label nuclei, enabling cell hashing and doublet detection. Overview of sci-Plex: Cells corresponding to different perturbations were lysed in wells, and their nuclei were labeled with well-specific "hash" oligos, followed by fixation, pooling, and sci-RNA-seq. [Figure 4C] Figure 1 shows that SqueePLEX uses polyadenylated single-stranded oligonucleotides to label nuclei, enabling cell hashing and doublet detection. Scatter plot showing the number of UMIs from single-cell transcriptomes derived from a mixture of hashed human HEK293T and mouse NIH3T3 cells. Points are colored based on hashed oligo assignment. [Figure 4D] Figure 1 shows that SqueePLEX uses polyadenylated single-stranded oligonucleotides to label nuclei, enabling cell hashing and doublet detection. Boxplot showing the number of mRNA UMIs recovered per cell for fresh versus frozen human and mouse cell lines. [Figure 4E] Figure 1 shows that SqueePLEX uses polyadenylated single-stranded oligonucleotides to label nuclei, enabling cell hashing and doublet detection. Scatter plot of overload experiment. Axes are the same as in (C). Identification has oligo collisions (red), which identifies cell collisions with high sensitivity.

[0067] [Figure 5A]We demonstrate that hashing with short polyadenylated single-stranded oligonucleotides enables stable, low-cost labeling of nuclei for sci-RNA-seq and subsequent doublet detection. Fluorescence microscopy images showing the lack of Alexa 647-conjugated oligo staining (right) in non-permeabilized H3-GFP+ NIH3T3 cells (left). [Figure 5B] We demonstrate that hashing with short polyadenylated single-stranded oligonucleotides allows stable, low-cost labeling of nuclei for sci-RNA-seq and subsequent doublet detection. Design of the polyadenylated hash oligo (top) and the indexed primer (bottom) used for reverse transcription. [Figure 5C] We demonstrate that hashing with short polyadenylated single-stranded oligonucleotides allows for stable, low-cost labeling of nuclei for sci-RNA-seq and subsequent doublet detection. Number of hashed UMIs detected per cell. Cells with fewer than 10 hashed UMIs (red line) were excluded from further analysis. [Figure 5D] We demonstrate that hashing with short polyadenylated single-stranded oligonucleotides allows for stable, low-cost labeling of nuclei for sci-RNA-seq and subsequent doublet detection. Distribution of enrichment ratios in cells. The enrichment ratio was calculated as the UMI count ratio of the most abundant to the second most abundant hashed oligo. An enrichment ratio cutoff of 15 (red line) was used to distinguish doublets from singlets. [Figure 5E] Figure 1 shows that hashing with short polyadenylated single-stranded oligonucleotides allows stable, low-cost labeling of nuclei for sci-RNA-seq and subsequent doublet detection. Boxplot of the number of cells recovered per well for each cell line. [Figure 5F]We demonstrate that hashing with short polyadenylated single-stranded oligonucleotides allows for stable, low-cost labeling of nuclei for sci-RNA-seq and subsequent doublet detection. The layout of the culture plate wells, with colors indicating the number of cells recovered, and outlines indicating the cell line. Note that although more NIH3T3 cells were recovered per well, similar numbers of cells were recovered across wells of each cell type. [Figure 5G] We show that hashing with short polyadenylated single-stranded oligonucleotides allows for stable, low-cost labeling of nuclei for sci-RNA-seq and subsequent doublet detection. UMI counts normalized by size factor, tabulated per gene on a logarithmic scale, recovered from sci-RNA-seq of fresh versus frozen preparations. The size factor is calculated as the logarithmic count observed in a single cell divided by the geometric mean of the logarithmic counts from all measured cells. The black line indicates y = x. The red line is a fit with the Pearson correlation shown. [Figure 5H] Log-scale boxplot of the number of hashed UMIs recovered from sci-RNA-seq of HEK293T (human) or NIH3T3 (mouse) cells from fresh versus frozen preparations. [Figure 5I] Figure 1 shows that hashing with short polyadenylated single-stranded oligonucleotides allows stable, low-cost labeling of nuclei for sci-RNA-seq and subsequent doublet detection. Theoretical (red bars) vs. observed (black dots are individual wells and blue bars are average) doublet rates as a function of the number of nuclei sorted onto the final plate during sci-RNA-seq. [Figure 5J] We demonstrate that hashing with short polyadenylated single-stranded oligonucleotides allows stable, low-cost labeling of nuclei for sci-RNA-seq and subsequent doublet detection. The barnyard plot in Figure 1E after removal of doublets detected by hashing. [Figure 5K] We show that hashing with short polyadenylated single-stranded oligonucleotides enables stable, low-cost labeling of nuclei for sci-RNA-seq and subsequent doublet detection. Log-scale boxplot of the number of RNA UMIs in singlet versus doublet cells, called based on the purity of hashed UMIs. Notably, these are "intra-species" doublets, i.e., human-human or mouse-mouse, that are not easily detected in traditional barnyard experiments.

[0068] [Figure 6A] This demonstrates that sci-Plex enables multiplexed chemical transcriptomics at single-cell resolution. Figure 1 shows the compounds and corresponding targets assayed within a pilot sci-Plex experiment. A549 lung adenocarcinoma cells were treated with either vehicle [dimethyl sulfoxide (DMSO) or ethanol] or one of four compounds (BMS345541, dexamethasone, Nutlin-3a, or SAHA). [Figure 6B] Showing that sci-Plex enables multiplexed chemical transcriptomics at single-cell resolution. UMAP embedding of chemically perturbed A549 cells stained by drug treatment. [Figure 6C] Showing that sci-Plex enables multiplexed chemical transcriptomics at single-cell resolution. UMAP embedding of chemically perturbed A549 cells faceted by treatment with cells colored by dose. [Figure 6D] Figure 1 shows that sci-Plex enables multiplexed chemical transcriptomics at single-cell resolution. (D and E) Expression of canonical (D) glucocorticoid receptor activating (ANGPTL4) and repressing (GDF15) target genes as a function of dexamethasone dose, or (E) p53 target genes as a function of Nutlin-3a dose. The y-axis indicates the percentage of cells with at least one read corresponding to the transcript. [Figure 6E]Figure 1 shows that sci-Plex enables multiplexed chemical transcriptomics at single-cell resolution. (D and E) Expression of canonical (D) glucocorticoid receptor activating (ANGPTL4) and repressing (GDF15) target genes as a function of dexamethasone dose, or (E) p53 target genes as a function of Nutlin-3a dose. The y-axis indicates the percentage of cells with at least one read corresponding to the transcript. [Figure 6F] Figure 1 shows that sci-Plex enables multiplexed chemical transcriptomics at single-cell resolution. Dose-response viability estimates for A549 cells treated with BMS345541, dexamethasone, Nutlin-3a, and SAHA based on the relative number of cells recovered at each dose.

[0069] [Figure 7A] We demonstrate that sci-Plex differentiates the transcriptional responses of A549 cells to four small molecules and recovers dose-response estimates similar to established assays. Experimental layout of A549 cells in a 96-well plate. Cells were treated for 24 hours in two 96-well plates with seven doses (or vehicle) arranged along each column. [Figure 7B] Figure 1 shows that sci-Plex differentiates the transcriptional responses of A549 cells to four small molecules and recovers dose-response estimates similar to established assays. B) Cells containing >30 hashed oligo UMIs and C) enrichment ratios >10 were retained. [Figure 7C] Figure 1 shows that sci-Plex differentiates the transcriptional responses of A549 cells to four small molecules and recovers dose-response estimates similar to established assays. B) Cells containing >30 hashed oligo UMIs and C) enrichment ratios >10 were retained. [Figure 7D] We demonstrate that sci-Plex differentiates the transcriptional responses of A549 cells to four small molecules and recovers dose-response estimates similar to established assays. Retained cells had a median hash UMI count of 78 and a median RNA UMI count of 4,681. [Figure 7E] We demonstrate that sci-Plex differentiates the transcriptional responses of A549 cells to four small molecules and recovers dose-response estimates similar to established assays. UMAP embeddings of chemically perturbed A549 cells are comparable to those in Figure 2B, except that cells are colored according to whether they were treated with vehicle or one of the four small molecules. [Figure 7F] We demonstrate that sci-Plex differentiates the transcriptional responses of A549 cells to four small molecules and recovers dose-response estimates similar to established assays. The UMAP embedding of chemically perturbed A549 cells is comparable to Figure 2B, except that cells are colored by clusters defined using the density peak algorithm in Monocle 3. [Figure 7G] We demonstrate that sci-Plex differentiates the transcriptional responses of A549 cells to four small molecules and recovers dose-response estimates similar to established assays. Ketron shows how pools of barcoded nuclei preserve relative cell numbers. [Figure 7H] Figure 1 shows that sci-Plex differentiates the transcriptional responses of A549 cells to four small molecules and recovers dose-response estimates similar to established assays. Viability estimates by counting the percentage of recovered hashed nuclei (gray) vs. CellTiter-Glo (red, n=6). [Figure 7I] Figure 1 shows that sci-Plex differentiates the transcriptional responses of A549 cells to four small molecules and recovers dose-response estimates similar to established assays. Scatter plots of estimated cell number (x-axis) and CellTiter-Glo viability estimates (y-axis) across all treatments and doses tested (Pearson correlation and chi-squared test).

[0070] [Figure 8A]Dose-dependent differentially expressed genes (DEGs) are shown to recover predicted transcriptional modules. Upset graphs display the intersection of dose-dependent DEGs between treatments (vertical bars) as well as the total number of dose-dependent DEGs per treatment (horizontal bars). Genes are defined as dose-dependent DEGs if a quasi-Poisson regression model relating expression in a specific cell to the dose of drug the cell received shows a significant dose effect (Wald test) after Benjamini-Hochberg correction (FDR < 0.05). See Methods for complete details on regression modeling. The four leftmost bars correspond to drug-specific dose-dependent DEGs, while the rightmost vertical bars correspond to dose-dependent DEGs shared by all four drugs. [Figure 8B] Figure 1 shows that dose-dependent differentially expressed genes (DEGs) recover predicted transcriptional modules. Gene set analysis (GSA) was performed on dose-dependent DEGs using the runGSA() function in the Piano package and the Hallmarks gene set from MSigDB (45). Heatmap colors indicate the value of the directional GSA enrichment statistic, with values ​​capped at either -10 or +10 for visualization.

[0071] [Figure 9A] This demonstrates that sci-Plex enables global transcriptional profiling of thousands of chemical perturbations in a single experiment. Schematic of a large-scale sci-Plex experiment (sciRNA-seq3). After 24 hours of treatment, a total of 188 small molecules were tested for their effects on A549, K562, and MCF7 human cell lines, each with four doses and biological replicates. Dose and drug plate position were varied between replicates, and a median of 100 to 200 cells was collected per condition. Colors distinguish cell lines, compound pathways, and doses. [Figure 9B]This demonstrates that sci-Plex enables global transcriptional profiling of thousands of chemical perturbations in a single experiment. UMAP embedding of A549, K562, and MCF7 cells from our screen. Each cell is colored according to the pathway targeted by the compound to which that particular cell was exposed. To facilitate visualization of significant molecular phenotypes, we added transparency for cells treated with compound or dose combinations that did not significantly alter the distribution of corresponding cells in UMAP space compared to vehicle controls (Fisher's exact test, FDR < 1%). [Figure 9C] This demonstrates that sci-Plex enables global transcriptional profiling of thousands of chemical perturbations in a single experiment. Viability estimates derived from hash-based counts of nuclei at each dose of selected compounds (bosutinib is highlighted in red text). Rows represent increasing compound doses from top to bottom, while columns represent individual compounds. The annotation bar at the top indicates the broad cellular activity targeted by each compound. [Figure 9D] Showing that sci-Plex enables global transcriptional profiling of thousands of chemical perturbations in a single experiment, UMAP embedding highlighted by treatment with the MEK inhibitor trametinib (red), an HSP90 inhibitor (purple), or vehicle control (gray). [Figure 9E] This demonstrates that sci-Plex enables global transcriptional profiling of thousands of chemical perturbations in a single experiment. HSP90AA1 expression levels in cells exposed to increasing doses of trametinib. The y-axis shows the percentage of cells with at least one read corresponding to the transcript.

[0072] [Figure 10A]This figure shows hash-based cell labeling in a large-scale sci-Plex experiment. The sci-Plex hashing design has 188 compounds. The experiment used 52 96-well plates, with each well marked by a combination of two oligos: one specific to a single 96-well culture plate and one specific to a well within that culture plate. [Figure 10B] We demonstrate hash-based cell labeling in a large-scale sci-Plex experiment. While this could theoretically be implemented with simply 96 wells of hash oligos, we instead used 768. This means that of the 39,936 possible pairings of plate and well hash oligos, only a small number (12.5%) of combinations were predicted as "legitimate," while most were unexpectedly "illegitimate." [Figure 10C] Hash-based cell labeling in a large-scale sci-Plex experiment. The observed pairings of plate and well hash oligos were strongly enriched for "correct" combinations. [Figure 10D] Scatter plot of HEK293T and NIH3T3 cells seeded in a single RT well of a large-scale sci-Plex experiment showing hash-based cell labeling in a large-scale sci-Plex experiment. [Figure 10E] Hash-based cell labeling in large-scale sci-Plex experiments. E-H) Hash UMI (Panels E and G) and enrichment ratio (Panels F and H) cutoffs used for well hash oligos (Panels E and F) and plate hash oligos (Panels G and H). The enrichment ratio cutoff corresponds to >5-fold enrichment. The hash UMI cutoff corresponds to >5. [Figure 10F] Hash-based cell labeling in large-scale sci-Plex experiments. E-H) Hash UMI (Panels E and G) and enrichment ratio (Panels F and H) cutoffs used for well hash oligos (Panels E and F) and plate hash oligos (Panels G and H). The enrichment ratio cutoff corresponds to >5-fold enrichment. The hash UMI cutoff corresponds to >5. [Figure 10G] Hash-based cell labeling in large-scale sci-Plex experiments. E-H) Hash UMI (Panels E and G) and enrichment ratio (Panels F and H) cutoffs used for well hash oligos (Panels E and F) and plate hash oligos (Panels G and H). The enrichment ratio cutoff corresponds to >5-fold enrichment. The hash UMI cutoff corresponds to >5. [Figure 10H] Hash-based cell labeling in large-scale sci-Plex experiments. E-H) Hash UMI (Panels E and G) and enrichment ratio (Panels F and H) cutoffs used for well hash oligos (Panels E and F) and plate hash oligos (Panels G and H). The enrichment ratio cutoff corresponds to >5-fold enrichment. The hash UMI cutoff corresponds to >5.

[0073] [Figure 11A] Quality control metrics for large-scale sci-Plex experiments are shown. Log-scale boxplots of the number of RNA UMIs for cells passing the hash and RNA UMI cutoff filters for each of the three cell lines. [Figure 11B] Figure 1 shows quality control metrics for a large-scale sci-Plex experiment. Correlation of size-factor normalized counts for genes between replicates for each of the three cell lines. The black line indicates y = x. The red line is a fit with the Pearson correlation shown. [Figure 11C] Figure 1 shows quality control metrics for large-scale sci-Plex experiments. Boxplots showing the number of vehicle cells recovered from each of the eight vehicle control wells within each replicate for A549, K562, and MCF7 cells.

[0074] [Figure 12A]Exposing cells to compounds alters their distribution across cell clusters. Heatmap showing the log-transformed ratio of cells treated with a specific drug relative to vehicle-control cells within each Louvain community. Columns correspond to clusters in PCA space (see Figures 13A-C), and rows correspond to compounds, annotated by pathway and target. Gray entries indicate compounds that are not significantly enriched or depleted relative to vehicle within the corresponding cluster (Fisher's exact test, FDR < 1%). [Figure 12B] Exposing cells to compounds alters their distribution across cell clusters. Heatmap showing the log-transformed ratio of cells treated with a specific drug relative to vehicle-control cells within each Louvain community. Columns correspond to clusters in PCA space (see Figures 13A-C), and rows correspond to compounds, annotated by pathway and target. Gray entries indicate compounds that are not significantly enriched or depleted relative to vehicle within the corresponding cluster (Fisher's exact test, FDR < 1%). [Figure 12C] Exposing cells to compounds alters their distribution across cell clusters. Heatmap showing the log-transformed ratio of cells treated with a specific drug relative to vehicle-control cells within each Louvain community. Columns correspond to clusters in PCA space (see Figures 13A-C), and rows correspond to compounds, annotated by pathway and target. Gray entries indicate compounds that are not significantly enriched or depleted relative to vehicle within the corresponding cluster (Fisher's exact test, FDR < 1%). [Figure 12D]Exposing cells to compounds alters their distribution across cell clusters. Heatmap showing the log-transformed ratio of cells treated with a specific drug relative to vehicle-control cells within each Louvain community. Columns correspond to clusters in PCA space (see Figures 13A-C), and rows correspond to compounds, annotated by pathway and target. Gray entries indicate compounds that are not significantly enriched or depleted relative to vehicle within the corresponding cluster (Fisher's exact test, FDR < 1%).

[0075] [Figure 13A] Showing that sci-Plex identifies pathway-specific enrichment of compounds across UMAP clusters. A-C) UMAP embeddings from Figure 3B colored by cell assignment to Louvain communities across PCA space for A549 (panel A), K562 (panel B), and MCF7 (panel C) cells. [Figure 13B] Showing that sci-Plex identifies pathway-specific enrichment of compounds across UMAP clusters. A-C) UMAP embeddings from Figure 3B colored by cell assignment to Louvain communities across PCA space for A549 (panel A), K562 (panel B), and MCF7 (panel C) cells. [Figure 13C] Showing that sci-Plex identifies pathway-specific enrichment of compounds across UMAP clusters. A-C) UMAP embeddings from Figure 3B colored by cell assignment to Louvain communities across PCA space for A549 (panel A), K562 (panel B), and MCF7 (panel C) cells. [Figure 13D]This shows that sci-Plex identifies pathway-specific enrichment of compounds across UMAP clusters. UMAP embedding of A549 cells from Figure 3B. Cells treated with the glucocorticoid receptor (GR) agonist triamcinolone acetonide are highlighted in green, while all other cells are highlighted in gray. These cells comprise the majority (95%) of cells in cluster 18 in panel A. [Figure 13E] Figure 1 shows that sci-Plex identifies pathway-specific enrichment of compounds across UMAP clusters. Percentage of A549 cells expressing the GR target genes ANGPTL4 and GDF15 as a function of increasing doses of the synthetic GR agonist triamcinolone acetonide. [Figure 13F] This demonstrates that sci-Plex identifies pathway-specific enrichment of compounds across UMAP clusters. F-H) UMAP embeddings of A549 cells stained by cells treated with various doses of epothilone A (F), epothilone B (G), or by proliferation index (H). Insets show magnified views of distinct foci induced upon treatment. The treatment with the highest number of cells within each bounding box is shown in panel H, with the cell count in parentheses. [Figure 13G] This demonstrates that sci-Plex identifies pathway-specific enrichment of compounds across UMAP clusters. F-H) UMAP embeddings of A549 cells stained by cells treated with various doses of epothilone A (F), epothilone B (G), or by proliferation index (H). Insets show magnified views of distinct foci induced upon treatment. The treatment with the highest number of cells within each bounding box is shown in panel H, with the cell count in parentheses. [Figure 13H]This demonstrates that sci-Plex identifies pathway-specific enrichment of compounds across UMAP clusters. F-H) UMAP embeddings of A549 cells stained by cells treated with various doses of epothilone A (F), epothilone B (G), or by proliferation index (H). Insets show magnified views of distinct foci induced upon treatment. The treatment with the highest number of cells within each bounding box is shown in panel H, with the cell count in parentheses.

[0076] [Figure 14] The number of dose-dependent differentially expressed genes detected per compound category is shown. Genes with significant dose-dependent differential expression (FDR<0.05) are grouped by cell line and colored by target pathway.

[0077] [Figure 15A] Correlation between "pseudobulk" sci-Plex and bulk-RNA-seq is shown. Log10 transcripts per million (TPM) of protein-coding genes measured by bulk RNA-seq (x-axis) and aggregate single-cell profiles normalized by the size factor of vehicle-treated cells from sci-Plex (y-axis). Results are shown for both A549 and K562 cells. The black line indicates the line y = x, and the blue line indicates a linear fit with the Pearson correlation shown. [Figure 15B] Correlation between "pseudobulk" sci-Plex and bulk-RNA-seq is shown. Scatter plot of selected compounds comparing statistically significant estimates obtained from linear models fit to single-cell data (x-axis) with estimates obtained from bulk RNA-seq using DESeq2 (y-axis). Black line indicates y = x. Blue line is fit with Pearson correlation shown.

[0078] [Figure 16A]The graph shows that the mean Z-scores from the L1000 assay correlate with the dose-dependent betas from sci-Plex. For a selected compound-cell line combination (trichostatin A in MCF7 cells), the mean Z-scores from the L1000 assay treated for 24 hours with each of the eight doses (y-axis) (11) versus the dose-dependent betas from the sci-Plex data are plotted. All genes that were part of the L1000 assay and had a significant dose-dependent effect by sci-Plex (p-value < 0.01) are shown. The lines are fits with Spearman correlations shown. [Figure 16B] Figure 1 shows that the moderate Z-scores from the L1000 assay correlate with the dose-dependent betas from sci-Plex. Figure 1 shows the Spearman correlation between significant sci-Plex-calculated dose-dependent betas and L1000 moderate Z-score values ​​from the LINCS L1000 data for genes measured at the highest dose in MCF7 cells. Compounds are presented grouped by the pathway they target. The red dot corresponds to fulvestrant. [Figure 16C] Showing that the moderate Z-score from the L1000 assay correlates with the dose-dependent beta from sci-Plex. Same as panel A, but for fulvestrant in MCF7 cells at the highest dose (10 μM). [Figure 16D] A moderate Z-score from the L1000 assay correlates with dose-dependent beta from sci-Plex. Same as panel B, but for A549 cells. Red dots correspond to triamcinolone acetonide. [Figure 16E] Showing that the moderate Z-score from the L1000 assay correlates with the dose-dependent beta from sci-Plex. Same as panel A, but for triamcinol acetonide in A549 cells at the highest dose (10 μM).

[0079] [Figure 17A]Single-cell measurements reveal changes in proliferation status in vehicle-treated cells and across doses of each drug. A-C) UMAP projections of A549 (A), K562 (B), and MCF7 (C) stained by proliferation index. High proliferation index indicates increased aggregate expression of transcripts that are markers of G1 / S or G2 / M phase (43). [Figure 17B] Single-cell measurements reveal changes in proliferation status in vehicle-treated cells and across doses of each drug. A-C) UMAP projections of A549 (A), K562 (B), and MCF7 (C) stained by proliferation index. High proliferation index indicates increased aggregate expression of transcripts that are markers of G1 / S or G2 / M phase (43). [Figure 17C] Single-cell measurements reveal changes in proliferation status in vehicle-treated cells and across doses of each drug. A-C) UMAP projections of A549 (A), K562 (B), and MCF7 (C) stained by proliferation index. High proliferation index indicates increased aggregate expression of transcripts that are markers of G1 / S or G2 / M phase (43). [Figure 17D] Figure 1 shows that single-cell measurements reveal changes in proliferation state in vehicle-treated cells and across doses of each drug. (DF) Density graphs of cell cycle distribution for compound-treated (blue fill) or vehicle-treated (red line) cells. The gray line indicates the cutoff used to distinguish between proliferating cells (above the cutoff) and non-proliferating cells (below the cutoff). [Figure 17E] Figure 1 shows that single-cell measurements reveal changes in proliferation state in vehicle-treated cells and across doses of each drug. (DF) Density graphs of cell cycle distribution for compound-treated (blue fill) or vehicle-treated (red line) cells. The gray line indicates the cutoff used to distinguish between proliferating cells (above the cutoff) and non-proliferating cells (below the cutoff). [Figure 17F]Figure 1 shows that single-cell measurements reveal changes in proliferation state in vehicle-treated cells and across doses of each drug. (DF) Density graphs of cell cycle distribution for compound-treated (blue fill) or vehicle-treated (red line) cells. The gray line indicates the cutoff used to distinguish between proliferating cells (above the cutoff) and non-proliferating cells (below the cutoff). [Figure 17G] Figure 1 shows that single-cell measurements reveal changes in proliferation status in vehicle-treated cells and across doses of each drug. GI) Percentage of cells designated as low-proliferative at each dose of each drug (x-axis) versus estimated median viability for that combination (y-axis). Each black dot corresponds to a cell treated with the same dose of a given drug. Red dots correspond to vehicle treatment. [Figure 17H] Figure 1 shows that single-cell measurements reveal changes in proliferation status in vehicle-treated cells and across doses of each drug. GI) Percentage of cells designated as low-proliferative at each dose of each drug (x-axis) versus estimated median viability for that combination (y-axis). Each black dot corresponds to a cell treated with the same dose of a given drug. Red dots correspond to vehicle treatment. [Figure 17I] Figure 1 shows that single-cell measurements reveal changes in proliferation status in vehicle-treated cells and across doses of each drug. GI) Percentage of cells designated as low-proliferative at each dose of each drug (x-axis) versus estimated median viability for that combination (y-axis). Each black dot corresponds to a cell treated with the same dose of a given drug. Red dots correspond to vehicle treatment. [Figure 17J] Figure 1 shows that single cell measurements reveal changes in proliferation status in vehicle-treated cells and across doses of each drug. Volcanograph showing the log2 fold change of genes that were significantly (q value < 0.01) differentially expressed between high and low fractions of vehicle-treated cells.

[0080] [Figure 18A]This demonstrates that single-cell measurements allow estimation of proliferation state and viability across drug-dose combinations. The heatmap shows estimates of relative proliferation rate, the percentage of cells exhibiting a low proliferation index, and estimated viability for each compound (rows) at each dose (columns). [Figure 18B] This demonstrates that single-cell measurements allow estimation of proliferation state and viability across drug-dose combinations. The heatmap shows estimates of relative proliferation rate, the percentage of cells exhibiting a low proliferation index, and estimated viability for each compound (rows) at each dose (columns).

[0081] [Figure 19A] sci-Plex demonstrates that it enables the analysis of proliferating and non-proliferating cell populations. Schematic showing how changes in cell state (top) and relative frequencies of subpopulations (bottom) appear identical when samples are subjected to aggregate measurements such as bulk RNA-seq. Adapted from reference (14). [Figure 19B] This shows that sci-Plex enables the analysis of proliferating and non-proliferating cell populations. B, C) Pearson correlation between dose-dependent effect sizes estimated from high and low proliferation index cells for each cell line (panel B) and drug class (panel C). [Figure 19C] This shows that sci-Plex enables the analysis of proliferating and non-proliferating cell populations. B, C) Pearson correlation between dose-dependent effect sizes estimated from high and low proliferation index cells for each cell line (panel B) and drug class (panel C). [Figure 19D]This figure demonstrates that sci-Plex enables the analysis of proliferating and nonproliferating cell populations. Gene-specific effect sizes estimated from high (βdh) versus low (βdl) proliferation index cells for four selected compounds. Effect sizes are expressed as log2-transformed fold changes relative to the intercept. Four classes of genes are shown: those significant only in high proliferation index cells (green), only in low proliferation index cells (purple), in both high and low cells with concordant effect estimates (red), and in both high and low cells with discordant effect estimates (blue). If |βdh - βdl| was less than 10 percent of 1 / 2 (|βdh| + βdl|), the drug had a concordant dose-dependent effect on gene h in high (βdh) and low (βdh) cells. The black line indicates y = x.

[0082] [Figure 20A] Figure 1 shows that the sci-Plex screen identifies reproducible viability and expression signatures across validation experiments and orthogonal datasets. Cell count viability estimates for K562 (red), A549 (blue), and MCF7 (green) cells exposed to vehicle or increasing doses of the Src / Abl inhibitor bosutinib (n = 6 culture replicates, Wilcoxon rank-sum test). For each cell line, cell counts were normalized to the mean cell count of vehicle control-treated cells. Error bars represent the standard error of the mean, n = 8. [Figure 20B] The sci-Plex screen identifies reproducible viability and expression signatures across validation experiments and orthogonal datasets. EC50 values ​​for cell lines derived from hematopoietic and lymphoid systems, lung, and breast tissues, for which viability estimates are available from the Cancer Cell Line Encyclopedia (CCLE), exposed to the Abl inhibitor AZD0530 (left panel) or nilotinib (right panel). [Figure 20C]The sci-Plex screen demonstrates that it identifies reproducible viability and expression signatures across validation experiments and orthogonal datasets. C-E) Top connectivity scores (a measure summarizing the similarity between transcriptional signatures induced by different drugs (11, 12)) of MEK and HSP inhibitors from the CMAP database for all cell lines (summary, panel C) or for individual A549 (panel D) and MCF7 (panel E) cells. A connectivity score cutoff of + / - 90 was applied as in (11). [Figure 20D] The sci-Plex screen demonstrates that it identifies reproducible viability and expression signatures across validation experiments and orthogonal datasets. C-E) Top connectivity scores (a measure summarizing the similarity between transcriptional signatures induced by different drugs (11, 12)) of MEK and HSP inhibitors from the CMAP database for all cell lines (summary, panel C) or for individual A549 (panel D) and MCF7 (panel E) cells. A connectivity score cutoff of + / - 90 was applied as in (11). [Figure 20E] The sci-Plex screen demonstrates that it identifies reproducible viability and expression signatures across validation experiments and orthogonal datasets. C-E) Top connectivity scores (a measure summarizing the similarity between transcriptional signatures induced by different drugs (11, 12)) of MEK and HSP inhibitors from the CMAP database for all cell lines (summary, panel C) or for individual A549 (panel D) and MCF7 (panel E) cells. A connectivity score cutoff of + / - 90 was applied as in (11).

[0083] [Figure 21] Correlation of compound-driven molecular signatures to A549 cells identified in the sci-Plex screen. The heatmap shows the Pearson correlation of beta coefficients across dose-dependent differentially expressed genes for all pairwise combinations of screened compounds. To aid visualization, the Pearson correlation was limited to 0.6.

[0084] [Figure 22] Figure 1 shows the correlation of compound-driven molecular signatures to K562 cells identified in the sci-Plex screen. The heatmap shows the Pearson correlation of beta coefficients across dose-dependent differentially expressed genes for all pairwise combinations of screened compounds. To aid visualization, the Pearson correlation was limited to 0.6.

[0085] [Figure 23] Correlation of compound-driven molecular signatures to MCF7 cells identified in the sci-Plex screen. The heatmap shows the Pearson correlation of beta coefficients across dose-dependent differentially expressed genes for all pairwise combinations of screened compounds. To aid visualization, Pearson correlation was limited to 0.6.

[0086] [Figure 24A] Figure 1 shows cluster plots of compound-driven molecular signature correlations. The cluster plots show Pearson correlations of beta coefficients across dose-dependent differentially expressed genes for all pairwise combinations of compounds screened in A549 (A), K562 (B), and MCF7 (C) cells. Compound names are colored by pathway of interest. [Figure 24B] Figure 1 shows cluster plots of compound-driven molecular signature correlations. The cluster plots show Pearson correlations of beta coefficients across dose-dependent differentially expressed genes for all pairwise combinations of compounds screened in A549 (A), K562 (B), and MCF7 (C) cells. Compound names are colored by pathway of interest. [Figure 24C]Figure 1 shows cluster plots of compound-driven molecular signature correlations. The cluster plots show Pearson correlations of beta coefficients across dose-dependent differentially expressed genes for all pairwise combinations of compounds screened in A549 (A), K562 (B), and MCF7 (C) cells. Compound names are colored by pathway of interest.

[0087] [Figure 25] Figure 1 shows the UMAP embedding of drugs based on their dose-dependent effect on the expression of each gene. Each drug was provided to UMAP as a vector of effect estimates (see Methods) for all genes. Point shape corresponds to cell type, and color corresponds to compound class.

[0088] [Figure 26A] Pairwise distances between PCA embeddings of drugs based on their dose-dependent effects are shown. A) Heatmap of pairwise distances between two cell types (columns) for a specific drug (rows) in PCA reduced-dimensional space. Hierarchical clustering is performed to visualize cell-type-specific responses to each drug. B) Inset of the highlighted portion of the heatmap. Pathway annotations are shown on the left. Specific compounds highlighted with red arrows are shown on the right (C-E) as UMAP embeddings. F) Trametinib-treated cell lines are highlighted to demonstrate colocalization of A549 and K562. Colored dots correspond to labeled compounds; all other drugs are shown in gray. Shapes encode the cell line in which each effect profile was captured (square: MCF7, triangle: K562, circle: A549). [Figure 26B]Pairwise distances between PCA embeddings of drugs based on their dose-dependent effects are shown. A) Heatmap of pairwise distances between two cell types (columns) for a specific drug (rows) in PCA reduced-dimensional space. Hierarchical clustering is performed to visualize cell-type-specific responses to each drug. B) Inset of the highlighted portion of the heatmap. Pathway annotations are shown on the left. Specific compounds highlighted with red arrows are shown on the right (C-E) as UMAP embeddings. F) Trametinib-treated cell lines are highlighted to demonstrate colocalization of A549 and K562. Colored dots correspond to labeled compounds; all other drugs are shown in gray. Shapes encode the cell line in which each effect profile was captured (square: MCF7, triangle: K562, circle: A549). [Figure 26C] Pairwise distances between PCA embeddings of drugs based on their dose-dependent effects are shown. A) Heatmap of pairwise distances between two cell types (columns) for a specific drug (rows) in PCA reduced-dimensional space. Hierarchical clustering is performed to visualize cell-type-specific responses to each drug. B) Inset of the highlighted portion of the heatmap. Pathway annotations are shown on the left. Specific compounds highlighted with red arrows are shown on the right (C-E) as UMAP embeddings. F) Trametinib-treated cell lines are highlighted to demonstrate colocalization of A549 and K562. Colored dots correspond to labeled compounds; all other drugs are shown in gray. Shapes encode the cell line in which each effect profile was captured (square: MCF7, triangle: K562, circle: A549). [Figure 26D]Pairwise distances between PCA embeddings of drugs based on their dose-dependent effects are shown. A) Heatmap of pairwise distances between two cell types (columns) for a specific drug (rows) in PCA reduced-dimensional space. Hierarchical clustering is performed to visualize cell-type-specific responses to each drug. B) Inset of the highlighted portion of the heatmap. Pathway annotations are shown on the left. Specific compounds highlighted with red arrows are shown on the right (C-E) as UMAP embeddings. F) Trametinib-treated cell lines are highlighted to demonstrate colocalization of A549 and K562. Colored dots correspond to labeled compounds; all other drugs are shown in gray. Shapes encode the cell line in which each effect profile was captured (square: MCF7, triangle: K562, circle: A549). [Figure 26E] Pairwise distances between PCA embeddings of drugs based on their dose-dependent effects are shown. A) Heatmap of pairwise distances between two cell types (columns) for a specific drug (rows) in PCA reduced-dimensional space. Hierarchical clustering is performed to visualize cell-type-specific responses to each drug. B) Inset of the highlighted portion of the heatmap. Pathway annotations are shown on the left. Specific compounds highlighted with red arrows are shown on the right (C-E) as UMAP embeddings. F) Trametinib-treated cell lines are highlighted to demonstrate colocalization of A549 and K562. Colored dots correspond to labeled compounds; all other drugs are shown in gray. Shapes encode the cell line in which each effect profile was captured (square: MCF7, triangle: K562, circle: A549). [Figure 26F]Pairwise distances between PCA embeddings of drugs based on their dose-dependent effects are shown. A) Heatmap of pairwise distances between two cell types (columns) for a specific drug (rows) in PCA reduced-dimensional space. Hierarchical clustering is performed to visualize cell-type-specific responses to each drug. B) Inset of the highlighted portion of the heatmap. Pathway annotations are shown on the left. Specific compounds highlighted with red arrows are shown on the right (C-E) as UMAP embeddings. F) Trametinib-treated cell lines are highlighted to demonstrate colocalization of A549 and K562. Colored dots correspond to labeled compounds; all other drugs are shown in gray. Shapes encode the cell line in which each effect profile was captured (square: MCF7, triangle: K562, circle: A549).

[0089] [Figure 27A] HDAC inhibitor trajectories capture cellular heterogeneity in drug response and biochemical affinity. MNN alignment and UMAP embedding of transcriptional profiles of cells treated with one of 17 HDAC inhibitors. Sham-treated roots are indicated by red dots. [Figure 27B] Figure 1 shows that HDAC inhibitor trajectories capture cellular heterogeneity in drug response and biochemical affinity. Ridge graphs showing the distribution of cells along a simulated dose are shown for three HDAC inhibitors with varying biochemical affinities. [Figure 27C] Figure 1 shows that HDAC inhibitor trajectories capture cellular heterogeneity in drug response and biochemical affinity. Relationship between TC50 and mean log10(IC50) from in vitro measurements. Asterisks indicate compounds with solubility <200 mM (in DMSO) that were not included in the fit.

[0090] [Figure 28A]This shows that cell types treated with HDAC inhibitors are aligned, allowing for joint pseudo-administration trajectory reconstruction. UMAP embedding highlighting the reconstructed pseudo-administration trajectories on the mutually nearest neighbor aligned HDAC inhibitor and vehicle-treated cells. The root node (red dot) is selected as the node in the main graph, with more than 50% of its nearest neighbors annotated as vehicle-treated cells. [Figure 28B] The distribution of each cell line within the embedding shows that cell types treated with HDAC inhibitors align and co-distribute, allowing for reconstruction of the pseudo-administration trajectory. [Figure 28C] Figure 1 shows that cell types treated with HDAC inhibitors align and allow for joint pseudo-dosing trajectory reconstruction. Bar graph showing the percentage of each pseudo-dosing bin occupied by cells treated with each dose. [Figure 28D] Figure 1 shows that cell types treated with HDAC inhibitors align and allow for joint pseudo-dosing trajectory reconstruction. Bar graphs showing the percentage of each pseudo-dosing bin occupied by cells treated with each compound. [Figure 28E] Figure 1 shows that cell types treated with HDAC inhibitors align and co-represent the proportion within each pseudo-dosing bin corresponding to each cell line, allowing for reconstruction of the pseudo-dosing trajectory.

[0091] [Figure 29A] Ridge graphs show the distribution of cells along the sham treatment for each HDAC inhibitor and dose combination of compound localized to the HDAC trajectory. [Figure 29B] Ridge graphs show the distribution of cells along the sham treatment for each HDAC inhibitor and dose combination of compound localized to the HDAC trajectory. [Figure 29C] Ridge graphs show the distribution of cells along the sham treatment for each HDAC inhibitor and dose combination of compound localized to the HDAC trajectory.

[0092] [Figure 30A]Representative brightfield images of A549 cells exposed to vehicle (A) or specific doses of the SIRT1 activator SRT2104 (B) or the HDAC inhibitor abexinostat (C) show contact inhibition of cell proliferation after 72 hours of drug exposure. Viability estimates determined by the number of cells recovered for each drug / dose combination normalized to the cell number in vehicle control wells. [Figure 30B] Representative brightfield images of A549 cells exposed to vehicle (A) or specific doses of the SIRT1 activator SRT2104 (B) or the HDAC inhibitor abexinostat (C) show contact inhibition of cell proliferation after 72 hours of drug exposure. Viability estimates determined by the number of cells recovered for each drug / dose combination normalized to the cell number in vehicle control wells. [Figure 30C] Representative brightfield images of A549 cells exposed to vehicle (A) or specific doses of the SIRT1 activator SRT2104 (B) or the HDAC inhibitor abexinostat (C) show contact inhibition of cell proliferation after 72 hours of drug exposure. Viability estimates determined by the number of cells recovered for each drug / dose combination normalized to the cell number in vehicle control wells.

[0093] [Figure 31A] Alignment of A549 cells at 24 and 72 hours post-treatment reveals time-dependent responses to various small molecules. (A-C) UMAP embedding of A549 cells at 24 and 72 hours post-treatment without correction for differences in viability and proliferation (A), after linear transformation of the data to account for changes in proliferation index and viability (B), and after mutual nearest-neighbor-based alignment of the linearly transformed data (C). Cells are colored according to the time point at which they were collected. [Figure 31B]Alignment of A549 cells at 24 and 72 hours post-treatment reveals time-dependent responses to various small molecules. (A-C) UMAP embedding of A549 cells at 24 and 72 hours post-treatment without correction for differences in viability and proliferation (A), after linear transformation of the data to account for changes in proliferation index and viability (B), and after mutual nearest-neighbor-based alignment of the linearly transformed data (C). Cells are colored according to the time point at which they were collected. [Figure 31C] Alignment of A549 cells at 24 and 72 hours post-treatment reveals time-dependent responses to various small molecules. (A-C) UMAP embedding of A549 cells at 24 and 72 hours post-treatment without correction for differences in viability and proliferation (A), after linear transformation of the data to account for changes in proliferation index and viability (B), and after mutual nearest-neighbor-based alignment of the linearly transformed data (C). Cells are colored according to the time point at which they were collected. [Figure 31D] Alignment of A549 cells at 24 and 72 hours post-treatment reveals time-dependent responses to diverse small molecules. (DF) UMAP embedding as in panels AC, with cells colored by the aggregate normalized expression scores of G1 / S marker genes. [Figure 31E] Alignment of A549 cells at 24 and 72 hours post-treatment reveals time-dependent responses to diverse small molecules. (DF) UMAP embedding as in panels AC, with cells colored by the aggregate normalized expression scores of G1 / S marker genes. [Figure 31F] Alignment of A549 cells at 24 and 72 hours post-treatment reveals time-dependent responses to diverse small molecules. (DF) UMAP embedding as in panels AC, with cells colored by the aggregate normalized expression scores of G1 / S marker genes. [Figure 31G]Alignment of A549 cells at 24 and 72 h post-treatment reveals time-dependent responses to diverse small molecules. (GI) UMAP embedding as in panels AC, where cells are colored by the aggregate normalized expression scores of G2 / M marker genes. [Figure 31H] Alignment of A549 cells at 24 and 72 h post-treatment reveals time-dependent responses to diverse small molecules. (GI) UMAP embedding as in panels AC, where cells are colored by the aggregate normalized expression scores of G2 / M marker genes. [Figure 31I] Alignment of A549 cells at 24 and 72 h post-treatment reveals time-dependent responses to diverse small molecules. (GI) UMAP embedding as in panels AC, where cells are colored by the aggregate normalized expression scores of G2 / M marker genes. [Figure 31J] Alignment of A549 cells at 24 and 72 hours post-treatment reveals time-dependent responses to various small molecules. (JL) UMAP embedding as in panels AC, with cells colored by proliferation index. [Figure 31K] Alignment of A549 cells at 24 and 72 hours post-treatment reveals time-dependent responses to various small molecules. (JL) UMAP embedding as in panels AC, with cells colored by proliferation index. [Figure 31L] Alignment of A549 cells at 24 and 72 hours post-treatment reveals time-dependent responses to various small molecules. (JL) UMAP embedding as in panels AC, with cells colored by proliferation index. [Figure 31M] Alignment of A549 cells at 24 and 72 hours post-treatment reveals time-dependent responses to various small molecules. (MO) UMAP embedding as in panels A-C, visualizing only cells treated with vehicle control. [Figure 31N]Alignment of A549 cells at 24 and 72 hours post-treatment reveals time-dependent responses to various small molecules. (MO) UMAP embedding as in panels A-C, visualizing only cells treated with vehicle control. [Figure 31O] Alignment of A549 cells at 24 and 72 hours post-treatment reveals time-dependent responses to various small molecules. (MO) UMAP embedding as in panels A-C, visualizing only cells treated with vehicle control. [Figure 31P] Figure 1 shows alignment of A549 cells at 24 and 72 hours post-treatment revealing time-dependent responses to various small molecules. UMAP embedding from panel C, where cells are colored according to the pathway targeted by the treatment they were exposed to. [Figure 31Q] Figure 1 shows alignment of A549 cells at 24 and 72 hours post-treatment revealing time-dependent responses to various small molecules. Percentage of cells disrupted by pathways of interest. Note that only a subset of 188 compounds across a limited number of pathways was tested at 72 hours. [Figure 31R] Figure 1 shows alignment of A549 cells at 24 and 72 hours post-treatment revealing time-dependent responses to various small molecules. Percentage of cells disrupted by targeted activity upon treatment with epigenetic modulating compounds. [Figure 31S] Figure 1 shows alignment of A549 cells at 24 and 72 hours post-treatment revealing time-dependent responses to various small molecules. Percentage of cells disrupted by HDAC compounds.

[0094] [Figure 32A]Bromodomain inhibition, sirtuin activation, and histone deacetylase inhibition induce distinct transcriptomic responses. (A-D) UMAP embeddings of MNN-aligned A549 cells 24 and 72 hours after treatment with the pan-HDAC inhibitors abexinostat (A) or belinotat (B), the bromodomain inhibitor JQ1 (C), and the SIRT1 activator SRT2104 (D). Cells are colored according to the dose each cell was exposed to. [Figure 32B] Bromodomain inhibition, sirtuin activation, and histone deacetylase inhibition induce distinct transcriptomic responses. (A-D) UMAP embeddings of MNN-aligned A549 cells 24 and 72 hours after treatment with the pan-HDAC inhibitors abexinostat (A) or belinotat (B), the bromodomain inhibitor JQ1 (C), and the SIRT1 activator SRT2104 (D). Cells are colored according to the dose each cell was exposed to. [Figure 32C] Bromodomain inhibition, sirtuin activation, and histone deacetylase inhibition induce distinct transcriptomic responses. (A-D) UMAP embeddings of MNN-aligned A549 cells 24 and 72 hours after treatment with the pan-HDAC inhibitors abexinostat (A) or belinotat (B), the bromodomain inhibitor JQ1 (C), and the SIRT1 activator SRT2104 (D). Cells are colored according to the dose each cell was exposed to. [Figure 32D] Bromodomain inhibition, sirtuin activation, and histone deacetylase inhibition induce distinct transcriptomic responses. (A-D) UMAP embeddings of MNN-aligned A549 cells 24 and 72 hours after treatment with the pan-HDAC inhibitors abexinostat (A) or belinotat (B), the bromodomain inhibitor JQ1 (C), and the SIRT1 activator SRT2104 (D). Cells are colored according to the dose each cell was exposed to.

[0095] [Figure 33A]Figure 1 shows that the heterogeneous response to most HDAC inhibitors does not appear to be caused by cellular asynchrony. Aligned UMAP embeddings of cells exposed to vehicle HDAC inhibitor for 24 or 72 hours. Cells are stained according to progression along with sham treatment. [Figure 33B] Figure 1. Aligned UMAP embeddings of cells exposed to vehicle (gray cells) or labeled HDAC inhibitors for 24 h (red cells) or 72 h (blue cells), showing that the heterogeneous response to most HDAC inhibitors does not appear to be caused by cellular asynchrony. [Figure 33C] This shows that the heterogeneous response to most HDAC inhibitors does not appear to be caused by cellular asynchrony. Ridge graph showing the density of A549 cells exposed to HDAC inhibitors along aligned pseudo-treatment trajectories. Results are displayed for eight HDAC inhibitors assayed for both 24 and 72 hours. Gray and color-filled lines represent cells exposed to inhibitors for 24 or 72 hours, respectively.

[0096] [Figure 34A] Figure 1 shows transcriptional trajectories of cells treated with HDAC inhibitors corresponding to in vitro IC50 measurements. For each compound and each cell line, a pseudo-dose-response curve was fitted using the drc R package. The mean position of each dose along the pseudo-dose trajectory was used as the response. Two illustrative examples are shown: belinostat (top) and trichostatin A (TSA) (bottom). The dotted vertical lines indicate the transcriptional EC50 (TC50) of each compound in each cell line. The gray shaded areas indicate the 95% confidence interval for each TC50 estimate. [Figure 34B]Figure 1 shows that the transcriptional trajectories of cells treated with HDAC inhibitors correspond to in vitro IC50 measurements. Plot showing the mean log10(IC50[M]) vs. log(TC50) of in vitro measured aggregates colored by solubility, supplied by Selleckchem Chemicals. Points marked as (*) were not used for the fit. [Figure 34C] Figure 1 shows the transcriptional trajectories of cells treated with HDAC inhibitors corresponding to in vitro IC50 measurements. Log(IC50[M]) vs. log(TC50) for each HDAC isoform. Points are colored according to the HDAC inhibitor used.

[0097] [Figure 35A] Linear models identifying pseudo-dose-dependent modules of growth and metabolism are shown. Bar graphs of the total number of significant dose-dependent and pseudo-dose-dependent DEGs (FDR<0.05). [Figure 35B] Figure 1 shows a linear model identifying pseudo-dose-dependent modules of proliferation and metabolism. Figure 2 shows an upset plot displaying the intersection of significant pseudo-dose-dependent DEGs between the three cell types. [Figure 35C] A linear model identifying pseudo-dosage-dependent modules of growth and metabolism is shown. A pseudo-dosage heatmap showing 4,308 genes that changed significantly as a function of pseudo-dosage. Each row corresponds to the predicted expression of a gene in three cell lines fitted by the model described in the "Differential Expression Analysis" section of the method. Genes (rows) were scaled and normalized within each cell line before combining the three matrices and performing hierarchical clustering. Clusters from hierarchical clustering were used as input to GSAhyper using the Hallmark gene set collection. Selected genes and gene sets characterizing each cluster are shown (right).

[0098] [Figure 36A]Figure 1 shows that HDAC inhibitor treatment arrests the cell cycle in all three cell lines. Percentage of cells expressing AURKA and CDKN1A RNA across sham treatment bins. Black bars indicate bootstrapped 95% confidence intervals. [Figure 36B] Figure 1 shows that HDAC inhibitor treatment arrests the cell cycle in all three cell lines. Boxplots showing the percentage of cells in the low-proliferation fraction at a given drug dose across sham dosing bins. [Figure 36C] Figure 1 shows that HDAC inhibitor treatment arrests the cell cycle in all three cell lines. DNA content analysis of the three cell lines upon treatment with DMSO (top) or 10 μM abexinostat (bottom). [Figure 36D] Figure 1 shows that HDAC inhibitor treatment arrests the cell cycle in all three cell lines. Quantification of flow cytometry data showing the number of cells in each DNA content category.

[0099] [Figure 37A] This shows that exposure to HDAC inhibitors results in the sequestration of acetate in the form of acetylated lysine. Quantification of flow cytometry measurements of total cellular acetylated lysine in A549 (left panel), MCF7 (middle panel), and K562 (right panel) cells exposed to 10 μM pracinostat, 10 μM abexinostat, or vehicle control. Error bars indicate the standard deviation of the mean (Wilcoxon rank sum test, n = 3 culture replicates, *p < 0.05, ***p < 0.005). Representative flow cytometry histogram of the experiment quantified in panel A. The blue shaded area and red line correspond to the DMSO vehicle control and 10 μM abexinostat, respectively. [Figure 37B]This shows that exposure to HDAC inhibitors results in the sequestration of acetate in the form of acetylated lysine. Quantification of flow cytometry measurements of total cellular acetylated lysine in A549 (left panel), MCF7 (middle panel), and K562 (right panel) cells exposed to 10 μM pracinostat, 10 μM abexinostat, or vehicle control. Error bars indicate the standard deviation of the mean (Wilcoxon rank sum test, n = 3 culture replicates, *p < 0.05, ***p < 0.005). Representative flow cytometry histogram of the experiment quantified in panel A. The blue shaded area and red line correspond to the DMSO vehicle control and 10 μM abexinostat, respectively.

[0100] [Figure 38A] Figure 1 shows a shared transcriptional response to HDAC inhibitors indicative of acetyl-CoA depletion. Row-centered and z-scaled gene expression heatmap showing pseudo-dose-dependent upregulation of genes involved in cellular carbon metabolism. [Figure 38B] Figure 1 shows the HDAC inhibitor shared transcriptional response indicative of acetyl-CoA depletion. Diagram of the role of genes from (A) in cytoplasmic acetyl-CoA regulation. Red circles indicate acetyl groups. Enzymes are shown in gray. Transporters are shown in green (FA, fatty acid; Ac-CoA, acetyl-CoA; C, citrate).

[0101] [Figure 39A]Figure 1 shows that replenishment of acetyl-CoA precursors is reduced while inhibition of enzymes that replenish the acetyl-CoA pool worsens, indicating progression along a pseudo-dosing trajectory of HDAC inhibitors. A-D) UMAP embedding of A549 (panels A and B) and MCF7 (panels C and D) single-cell transcriptomes after exposure to the HDAC inhibitors pracinostat or abexinostat in the presence or absence of inhibitors of acetyl-CoA precursors or enzymes that replenish the acetyl-CoA pool. UMAPs were constructed from cells from all conditions in the experiment. Cells are colored by pseudo-dosing bin (panels A and C) or dose (panels B and D). E) Venn diagram of the overlap of differentially expressed genes across trajectories between the original HDACi trajectory versus the A549 or MCF7 HDACi trajectories from this new experiment. F, H) Boxplots of sham-dosed estimates for the selection condition of cells exposed to 1 or 10 μM pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel H) or MCF7 (panel L) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. G, I) Boxplots of sham-dosed estimates for the selection condition of cells exposed to vehicle and pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel I) or MCF7 (panel M) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. J, L) Heatmap showing the percentage of cells per sham-dosed bin exposed to various acetyl-coA precursors for A549 (F) or MCF7 (J) cells exposed to pranostin. K, M) Heatmap showing the percentage of cells per mock-dosed bin for cells exposed to various inhibitors targeting enzymes that replenish the acetyl-coA pool in A549 (panel G) and MCF7 (panel K) cells exposed to pracinostat. [Figure 39B]Figure 1 shows that replenishment of acetyl-CoA precursors is reduced while inhibition of enzymes that replenish the acetyl-CoA pool worsens, indicating progression along a pseudo-dosing trajectory of HDAC inhibitors. A-D) UMAP embedding of A549 (panels A and B) and MCF7 (panels C and D) single-cell transcriptomes after exposure to the HDAC inhibitors pracinostat or abexinostat in the presence or absence of inhibitors of acetyl-CoA precursors or enzymes that replenish the acetyl-CoA pool. UMAPs were constructed from cells from all conditions in the experiment. Cells are colored by pseudo-dosing bin (panels A and C) or dose (panels B and D). E) Venn diagram of the overlap of differentially expressed genes across trajectories between the original HDACi trajectory versus the A549 or MCF7 HDACi trajectories from this new experiment. F, H) Boxplots of sham-dosed estimates for the selection condition of cells exposed to 1 or 10 μM pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel H) or MCF7 (panel L) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. G, I) Boxplots of sham-dosed estimates for the selection condition of cells exposed to vehicle and pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel I) or MCF7 (panel M) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. J, L) Heatmap showing the percentage of cells per sham-dosed bin exposed to various acetyl-coA precursors for A549 (F) or MCF7 (J) cells exposed to pranostin. K, M) Heatmap showing the percentage of cells per mock-dosed bin for cells exposed to various inhibitors targeting enzymes that replenish the acetyl-coA pool in A549 (panel G) and MCF7 (panel K) cells exposed to pracinostat. [Figure 39C]Figure 1 shows that replenishment of acetyl-CoA precursors is reduced while inhibition of enzymes that replenish the acetyl-CoA pool worsens, indicating progression along a pseudo-dosing trajectory of HDAC inhibitors. A-D) UMAP embedding of A549 (panels A and B) and MCF7 (panels C and D) single-cell transcriptomes after exposure to the HDAC inhibitors pracinostat or abexinostat in the presence or absence of inhibitors of acetyl-CoA precursors or enzymes that replenish the acetyl-CoA pool. UMAPs were constructed from cells from all conditions in the experiment. Cells are colored by pseudo-dosing bin (panels A and C) or dose (panels B and D). E) Venn diagram of the overlap of differentially expressed genes across trajectories between the original HDACi trajectory versus the A549 or MCF7 HDACi trajectories from this new experiment. F, H) Boxplots of sham-dosed estimates for the selection condition of cells exposed to 1 or 10 μM pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel H) or MCF7 (panel L) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. G, I) Boxplots of sham-dosed estimates for the selection condition of cells exposed to vehicle and pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel I) or MCF7 (panel M) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. J, L) Heatmap showing the percentage of cells per sham-dosed bin exposed to various acetyl-coA precursors for A549 (F) or MCF7 (J) cells exposed to pranostin. K, M) Heatmap showing the percentage of cells per mock-dosed bin for cells exposed to various inhibitors targeting enzymes that replenish the acetyl-coA pool in A549 (panel G) and MCF7 (panel K) cells exposed to pracinostat. [Figure 39D]Figure 1 shows that replenishment of acetyl-CoA precursors is reduced while inhibition of enzymes that replenish the acetyl-CoA pool worsens, indicating progression along a pseudo-dosing trajectory of HDAC inhibitors. A-D) UMAP embedding of A549 (panels A and B) and MCF7 (panels C and D) single-cell transcriptomes after exposure to the HDAC inhibitors pracinostat or abexinostat in the presence or absence of inhibitors of acetyl-CoA precursors or enzymes that replenish the acetyl-CoA pool. UMAPs were constructed from cells from all conditions in the experiment. Cells are colored by pseudo-dosing bin (panels A and C) or dose (panels B and D). E) Venn diagram of the overlap of differentially expressed genes across trajectories between the original HDACi trajectory versus the A549 or MCF7 HDACi trajectories from this new experiment. F, H) Boxplots of sham-dosed estimates for the selection condition of cells exposed to 1 or 10 μM pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel H) or MCF7 (panel L) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. G, I) Boxplots of sham-dosed estimates for the selection condition of cells exposed to vehicle and pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel I) or MCF7 (panel M) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. J, L) Heatmap showing the percentage of cells per sham-dosed bin exposed to various acetyl-coA precursors for A549 (F) or MCF7 (J) cells exposed to pranostin. K, M) Heatmap showing the percentage of cells per mock-dosed bin for cells exposed to various inhibitors targeting enzymes that replenish the acetyl-coA pool in A549 (panel G) and MCF7 (panel K) cells exposed to pracinostat. [Figure 39E]Figure 1 shows that replenishment of acetyl-CoA precursors is reduced while inhibition of enzymes that replenish the acetyl-CoA pool worsens, indicating progression along a pseudo-dosing trajectory of HDAC inhibitors. A-D) UMAP embedding of A549 (panels A and B) and MCF7 (panels C and D) single-cell transcriptomes after exposure to the HDAC inhibitors pracinostat or abexinostat in the presence or absence of inhibitors of acetyl-CoA precursors or enzymes that replenish the acetyl-CoA pool. UMAPs were constructed from cells from all conditions in the experiment. Cells are colored by pseudo-dosing bin (panels A and C) or dose (panels B and D). E) Venn diagram of the overlap of differentially expressed genes across trajectories between the original HDACi trajectory versus the A549 or MCF7 HDACi trajectories from this new experiment. F, H) Boxplots of sham-dosed estimates for the selection condition of cells exposed to 1 or 10 μM pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel H) or MCF7 (panel L) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. G, I) Boxplots of sham-dosed estimates for the selection condition of cells exposed to vehicle and pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel I) or MCF7 (panel M) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. J, L) Heatmap showing the percentage of cells per sham-dosed bin exposed to various acetyl-coA precursors for A549 (F) or MCF7 (J) cells exposed to pranostin. K, M) Heatmap showing the percentage of cells per mock-dosed bin for cells exposed to various inhibitors targeting enzymes that replenish the acetyl-coA pool in A549 (panel G) and MCF7 (panel K) cells exposed to pracinostat. [Figure 39F]Figure 1 shows that replenishment of acetyl-CoA precursors is reduced while inhibition of enzymes that replenish the acetyl-CoA pool worsens, indicating progression along a pseudo-dosing trajectory of HDAC inhibitors. A-D) UMAP embedding of A549 (panels A and B) and MCF7 (panels C and D) single-cell transcriptomes after exposure to the HDAC inhibitors pracinostat or abexinostat in the presence or absence of inhibitors of acetyl-CoA precursors or enzymes that replenish the acetyl-CoA pool. UMAPs were constructed from cells from all conditions in the experiment. Cells are colored by pseudo-dosing bin (panels A and C) or dose (panels B and D). E) Venn diagram of the overlap of differentially expressed genes across trajectories between the original HDACi trajectory versus the A549 or MCF7 HDACi trajectories from this new experiment. F, H) Boxplots of sham-dosed estimates for the selection condition of cells exposed to 1 or 10 μM pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel H) or MCF7 (panel L) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. G, I) Boxplots of sham-dosed estimates for the selection condition of cells exposed to vehicle and pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel I) or MCF7 (panel M) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. J, L) Heatmap showing the percentage of cells per sham-dosed bin exposed to various acetyl-coA precursors for A549 (F) or MCF7 (J) cells exposed to pranostin. K, M) Heatmap showing the percentage of cells per mock-dosed bin for cells exposed to various inhibitors targeting enzymes that replenish the acetyl-coA pool in A549 (panel G) and MCF7 (panel K) cells exposed to pracinostat. [Figure 39G]Figure 1 shows that replenishment of acetyl-CoA precursors is reduced while inhibition of enzymes that replenish the acetyl-CoA pool worsens, indicating progression along a pseudo-dosing trajectory of HDAC inhibitors. A-D) UMAP embedding of A549 (panels A and B) and MCF7 (panels C and D) single-cell transcriptomes after exposure to the HDAC inhibitors pracinostat or abexinostat in the presence or absence of inhibitors of acetyl-CoA precursors or enzymes that replenish the acetyl-CoA pool. UMAPs were constructed from cells from all conditions in the experiment. Cells are colored by pseudo-dosing bin (panels A and C) or dose (panels B and D). E) Venn diagram of the overlap of differentially expressed genes across trajectories between the original HDACi trajectory versus the A549 or MCF7 HDACi trajectories from this new experiment. F, H) Boxplots of sham-dosed estimates for the selection condition of cells exposed to 1 or 10 μM pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel H) or MCF7 (panel L) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. G, I) Boxplots of sham-dosed estimates for the selection condition of cells exposed to vehicle and pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel I) or MCF7 (panel M) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. J, L) Heatmap showing the percentage of cells per sham-dosed bin exposed to various acetyl-coA precursors for A549 (F) or MCF7 (J) cells exposed to pranostin. K, M) Heatmap showing the percentage of cells per mock-dosed bin for cells exposed to various inhibitors targeting enzymes that replenish the acetyl-coA pool in A549 (panel G) and MCF7 (panel K) cells exposed to pracinostat. [Figure 39H]Figure 1 shows that replenishment of acetyl-CoA precursors is reduced while inhibition of enzymes that replenish the acetyl-CoA pool worsens, indicating progression along a pseudo-dosing trajectory of HDAC inhibitors. A-D) UMAP embedding of A549 (panels A and B) and MCF7 (panels C and D) single-cell transcriptomes after exposure to the HDAC inhibitors pracinostat or abexinostat in the presence or absence of inhibitors of acetyl-CoA precursors or enzymes that replenish the acetyl-CoA pool. UMAPs were constructed from cells from all conditions in the experiment. Cells are colored by pseudo-dosing bin (panels A and C) or dose (panels B and D). E) Venn diagram of the overlap of differentially expressed genes across trajectories between the original HDACi trajectory versus the A549 or MCF7 HDACi trajectories from this new experiment. F, H) Boxplots of sham-dosed estimates for the selection condition of cells exposed to 1 or 10 μM pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel H) or MCF7 (panel L) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. G, I) Boxplots of sham-dosed estimates for the selection condition of cells exposed to vehicle and pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel I) or MCF7 (panel M) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. J, L) Heatmap showing the percentage of cells per sham-dosed bin exposed to various acetyl-coA precursors for A549 (F) or MCF7 (J) cells exposed to pranostin. K, M) Heatmap showing the percentage of cells per mock-dosed bin for cells exposed to various inhibitors targeting enzymes that replenish the acetyl-coA pool in A549 (panel G) and MCF7 (panel K) cells exposed to pracinostat. [Figure 39I]Figure 1 shows that replenishment of acetyl-CoA precursors is reduced while inhibition of enzymes that replenish the acetyl-CoA pool worsens, indicating progression along a pseudo-dosing trajectory of HDAC inhibitors. A-D) UMAP embedding of A549 (panels A and B) and MCF7 (panels C and D) single-cell transcriptomes after exposure to the HDAC inhibitors pracinostat or abexinostat in the presence or absence of inhibitors of acetyl-CoA precursors or enzymes that replenish the acetyl-CoA pool. UMAPs were constructed from cells from all conditions in the experiment. Cells are colored by pseudo-dosing bin (panels A and C) or dose (panels B and D). E) Venn diagram of the overlap of differentially expressed genes across trajectories between the original HDACi trajectory versus the A549 or MCF7 HDACi trajectories from this new experiment. F, H) Boxplots of sham-dosed estimates for the selection condition of cells exposed to 1 or 10 μM pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel H) or MCF7 (panel L) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. G, I) Boxplots of sham-dosed estimates for the selection condition of cells exposed to vehicle and pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel I) or MCF7 (panel M) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. J, L) Heatmap showing the percentage of cells per sham-dosed bin exposed to various acetyl-coA precursors for A549 (F) or MCF7 (J) cells exposed to pranostin. K, M) Heatmap showing the percentage of cells per mock-dosed bin for cells exposed to various inhibitors targeting enzymes that replenish the acetyl-coA pool in A549 (panel G) and MCF7 (panel K) cells exposed to pracinostat. [Figure 39J]Figure 1 shows that replenishment of acetyl-CoA precursors is reduced while inhibition of enzymes that replenish the acetyl-CoA pool worsens, indicating progression along a pseudo-dosing trajectory of HDAC inhibitors. A-D) UMAP embedding of A549 (panels A and B) and MCF7 (panels C and D) single-cell transcriptomes after exposure to the HDAC inhibitors pracinostat or abexinostat in the presence or absence of inhibitors of acetyl-CoA precursors or enzymes that replenish the acetyl-CoA pool. UMAPs were constructed from cells from all conditions in the experiment. Cells are colored by pseudo-dosing bin (panels A and C) or dose (panels B and D). E) Venn diagram of the overlap of differentially expressed genes across trajectories between the original HDACi trajectory versus the A549 or MCF7 HDACi trajectories from this new experiment. F, H) Boxplots of sham-dosed estimates for the selection condition of cells exposed to 1 or 10 μM pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel H) or MCF7 (panel L) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. G, I) Boxplots of sham-dosed estimates for the selection condition of cells exposed to vehicle and pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel I) or MCF7 (panel M) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. J, L) Heatmap showing the percentage of cells per sham-dosed bin exposed to various acetyl-coA precursors for A549 (F) or MCF7 (J) cells exposed to pranostin. K, M) Heatmap showing the percentage of cells per mock-dosed bin for cells exposed to various inhibitors targeting enzymes that replenish the acetyl-coA pool in A549 (panel G) and MCF7 (panel K) cells exposed to pracinostat. [Figure 39K]Figure 1 shows that replenishment of acetyl-CoA precursors is reduced while inhibition of enzymes that replenish the acetyl-CoA pool worsens, indicating progression along a pseudo-dosing trajectory of HDAC inhibitors. A-D) UMAP embedding of A549 (panels A and B) and MCF7 (panels C and D) single-cell transcriptomes after exposure to the HDAC inhibitors pracinostat or abexinostat in the presence or absence of inhibitors of acetyl-CoA precursors or enzymes that replenish the acetyl-CoA pool. UMAPs were constructed from cells from all conditions in the experiment. Cells are colored by pseudo-dosing bin (panels A and C) or dose (panels B and D). E) Venn diagram of the overlap of differentially expressed genes across trajectories between the original HDACi trajectory versus the A549 or MCF7 HDACi trajectories from this new experiment. F, H) Boxplots of sham-dosed estimates for the selection condition of cells exposed to 1 or 10 μM pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel H) or MCF7 (panel L) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. G, I) Boxplots of sham-dosed estimates for the selection condition of cells exposed to vehicle and pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel I) or MCF7 (panel M) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. J, L) Heatmap showing the percentage of cells per sham-dosed bin exposed to various acetyl-coA precursors for A549 (F) or MCF7 (J) cells exposed to pranostin. K, M) Heatmap showing the percentage of cells per mock-dosed bin for cells exposed to various inhibitors targeting enzymes that replenish the acetyl-coA pool in A549 (panel G) and MCF7 (panel K) cells exposed to pracinostat. [Figure 39L]Figure 1 shows that replenishment of acetyl-CoA precursors is reduced while inhibition of enzymes that replenish the acetyl-CoA pool worsens, indicating progression along a pseudo-dosing trajectory of HDAC inhibitors. A-D) UMAP embedding of A549 (panels A and B) and MCF7 (panels C and D) single-cell transcriptomes after exposure to the HDAC inhibitors pracinostat or abexinostat in the presence or absence of inhibitors of acetyl-CoA precursors or enzymes that replenish the acetyl-CoA pool. UMAPs were constructed from cells from all conditions in the experiment. Cells are colored by pseudo-dosing bin (panels A and C) or dose (panels B and D). E) Venn diagram of the overlap of differentially expressed genes across trajectories between the original HDACi trajectory versus the A549 or MCF7 HDACi trajectories from this new experiment. F, H) Boxplots of sham-dosed estimates for the selection condition of cells exposed to 1 or 10 μM pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel H) or MCF7 (panel L) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. G, I) Boxplots of sham-dosed estimates for the selection condition of cells exposed to vehicle and pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel I) or MCF7 (panel M) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. J, L) Heatmap showing the percentage of cells per sham-dosed bin exposed to various acetyl-coA precursors for A549 (F) or MCF7 (J) cells exposed to pranostin. K, M) Heatmap showing the percentage of cells per mock-dosed bin for cells exposed to various inhibitors targeting enzymes that replenish the acetyl-coA pool in A549 (panel G) and MCF7 (panel K) cells exposed to pracinostat. [Figure 39M]Figure 1 shows that replenishment of acetyl-CoA precursors is reduced while inhibition of enzymes that replenish the acetyl-CoA pool worsens, indicating progression along a pseudo-dosing trajectory of HDAC inhibitors. A-D) UMAP embedding of A549 (panels A and B) and MCF7 (panels C and D) single-cell transcriptomes after exposure to the HDAC inhibitors pracinostat or abexinostat in the presence or absence of inhibitors of acetyl-CoA precursors or enzymes that replenish the acetyl-CoA pool. UMAPs were constructed from cells from all conditions in the experiment. Cells are colored by pseudo-dosing bin (panels A and C) or dose (panels B and D). E) Venn diagram of the overlap of differentially expressed genes across trajectories between the original HDACi trajectory versus the A549 or MCF7 HDACi trajectories from this new experiment. F, H) Boxplots of sham-dosed estimates for the selection condition of cells exposed to 1 or 10 μM pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel H) or MCF7 (panel L) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. G, I) Boxplots of sham-dosed estimates for the selection condition of cells exposed to vehicle and pranostin with or without co-treatment with acetyl-coA precursors for A549 (panel I) or MCF7 (panel M) cells. Values ​​are normalized to vehicle-treated cells. Wilcoxon rank-sum test. J, L) Heatmap showing the percentage of cells per sham-dosed bin exposed to various acetyl-coA precursors for A549 (F) or MCF7 (J) cells exposed to pranostin. K, M) Heatmap showing the percentage of cells per mock-dosed bin for cells exposed to various inhibitors targeting enzymes that replenish the acetyl-coA pool in A549 (panel G) and MCF7 (panel K) cells exposed to pracinostat.

[0102] [Figure 40A]Figure 1 shows the correlation of effect sizes between differentially expressed genes after HDAC inhibition from the original screen versus the new experiment. A-B) Correlation of effect size estimates (beta coefficients) for differentially expressed genes between vehicle control and 10 μM abexinostat (Panel A) or 10 μM pracinostat (Panel B) in A549 cells. C-D) Correlation of effect size estimates (beta coefficients) for differentially expressed genes between vehicle control and 10 μM abexinostat (Panel C) or 10 μM pracinostat (Panel D) in MCF7 cells. The X-axis corresponds to the large-scale sci-Plex experiment. The Y-axis corresponds to the targeted follow-up sci-Plex experiment. [Figure 40B] Figure 1 shows the correlation of effect sizes between differentially expressed genes after HDAC inhibition from the original screen versus the new experiment. A-B) Correlation of effect size estimates (beta coefficients) for differentially expressed genes between vehicle control and 10 μM abexinostat (Panel A) or 10 μM pracinostat (Panel B) in A549 cells. C-D) Correlation of effect size estimates (beta coefficients) for differentially expressed genes between vehicle control and 10 μM abexinostat (Panel C) or 10 μM pracinostat (Panel D) in MCF7 cells. The X-axis corresponds to the large-scale sci-Plex experiment. The Y-axis corresponds to the targeted follow-up sci-Plex experiment. [Figure 40C] Figure 1 shows the correlation of effect sizes between differentially expressed genes after HDAC inhibition from the original screen versus the new experiment. A-B) Correlation of effect size estimates (beta coefficients) for differentially expressed genes between vehicle control and 10 μM abexinostat (Panel A) or 10 μM pracinostat (Panel B) in A549 cells. C-D) Correlation of effect size estimates (beta coefficients) for differentially expressed genes between vehicle control and 10 μM abexinostat (Panel C) or 10 μM pracinostat (Panel D) in MCF7 cells. The X-axis corresponds to the large-scale sci-Plex experiment. The Y-axis corresponds to the targeted follow-up sci-Plex experiment. [Figure 40D]Figure 1 shows the correlation of effect sizes between differentially expressed genes after HDAC inhibition from the original screen versus the new experiment. A-B) Correlation of effect size estimates (beta coefficients) for differentially expressed genes between vehicle control and 10 μM abexinostat (Panel A) or 10 μM pracinostat (Panel B) in A549 cells. C-D) Correlation of effect size estimates (beta coefficients) for differentially expressed genes between vehicle control and 10 μM abexinostat (Panel C) or 10 μM pracinostat (Panel D) in MCF7 cells. The X-axis corresponds to the large-scale sci-Plex experiment. The Y-axis corresponds to the targeted follow-up sci-Plex experiment.

[0103] [Figure 41A] This demonstrates multiplexed single-cell ATAC-seq with combinatorial indexing, co-assaying nuclear labeling (hashing) and accessible chromatin. Cells from individual samples are grown and processed in separate wells. Within the processing wells, cells are lysed and nuclei are isolated. Well-specific single-stranded DNA oligo labels are added to each well, immobilizing the label within the nuclei. Nuclei from all samples are then pooled for downstream combinatorial indexing steps. [Figure 41B] Multiplexed single-cell ATAC-seq using combinatorial indexing methods co-assays nuclear labels (hashes) and accessible chromatin. Schematic of how combinatorial indexing adds the same molecular index to labels and accessible DNA in the nucleus. [Figure 41C] We demonstrate that multiplexed single-cell ATAC-seq, using a combinatorial indexing method, co-assays nuclear labels (hashed) and accessible chromatin. The recovered labels successfully distinguish cells from samples containing only mouse (NIH-3T3) or human (A549) cells in a barnyard mixing experiment. [Figure 41D]Multiplexed single-cell ATAC-seq using combinatorial indexing methods is shown to co-assay nuclear labels (hashes) and accessible chromatin. Nuclei containing multiple labels represent doublets. [Figure 41E] Multiplexed single-cell ATAC-seq using a combinatorial indexing method co-assays nuclear labeling (hashed) and accessible chromatin. Distribution of the number of labeled UMIs recovered per cell. The red line represents the cutoff requirement for including cells in downstream analyses. [Figure 41F] Multiplexed single-cell ATAC-seq using combinatorial indexing methods to co-assay nuclear labels (hashed) and accessible chromatin. Distribution of label enrichment ratios per cell. The enrichment ratio of a cell reflects the number of the most abundant label divided by the number of the second most abundant label in the nucleus. Cells below the red line (enrichment ratio <5) are referred to as doublets.

[0104] [Figure 42A] Figure 1 shows the labeling strategy that allows pairing of chromatin profiles from single cells to treatment groups. Experimental design. [Figure 42B] This figure shows a labeling strategy that allows pairing of chromatin profiles from single cells to treatment groups. Uniform manifold approximation and projection (UMAP) representations of all recovered cells stained by treatment, as determined by labeling. Note that the distance between two cells indicates the degree of difference in accessible chromatin landscapes. [Figure 42C] We demonstrate a labeling strategy that allows pairing of chromatin profiles from single cells to treatment groups. Dose-response curves for each drug are fitted to the number of cells recovered from each treatment dose. [Figure 42D] Figure 1 shows the labeling strategy that allows pairing of chromatin profiles from single cells to treatment groups. UMAP representation of cells from each drug treatment group, colored by the dose determined by labeling. [Figure 42E]Labeling strategy that allows pairing of chromatin profiles from single cells to treatment groups. Browser tracks showing pseudo-bulk ATAC profiles from cells treated with increasing doses of SAHA.

[0105] [Figure 43A] This demonstrates that hashed oligo ladders can be captured by nuclei and serve as external standards for sci-RNA-seq experiments. Experimental overview of the hashed ladder method: Nuclei are isolated from cells, fixed with hashed oligo ladders, and processed for sci-RNA-seq. [Figure 43B] Boxplot of hashed oligo UMI counts per cell, where each hashed oligo was spiked in at different amounts, showing that the hashed oligo ladder was captured by the nucleus and served as an external standard for sci-RNA-seq experiments. [Figure 43C] Figure 1 shows that the hash oligo ladder is captured by the nucleus and serves as an external standard for sci-RNA-seq experiments. Scatter plot of expected and observed hash ladder UMI counts, showing cells with low (left) and high (right) hash capture efficiency.

[0106] [Figure 44A] This demonstrates that hash ladders extend the ability to detect the overall decrease in transcript levels caused by flavopiridol. Experimental Overview: HEK293T cells were treated with flavopiridol for various periods and then labeled with a hash oligo ladder and additional hash oligos for multiplexing prior to sci-RNA-seq preparation. [Figure 44B] Boxplot showing total RNA UMI counts for cells treated with flavopiridol at different time points, demonstrating that hash ladders extend the ability to detect the overall decrease in transcript levels caused by flavopiridol. [Figure 44C]Figure 1 shows that hash ladders extend the ability to detect the overall decrease in transcript levels caused by flavopiridol. Bar graph showing the number of genes differentially expressed in response to flavopiridol using the conventional hash ladder normalization approach. [Figure 44D] Figure 1 shows that hash ladders extend the ability to detect the overall decrease in transcript levels caused by flavopralidol. Violin diagram showing the ratio of effect size estimates for common differentially expressed genes calculated with hash ladders versus conventional normalization.

[0107] Schematic diagrams are not necessarily to scale. Like numbers used in the figures refer to like components, steps, etc. However, it will be understood that the use of a number to refer to a component in a given figure is not intended to limit the component in another figure labeled with the same number. Furthermore, the use of different numbers to refer to a component is not intended to indicate that the differently numbered component may not be the same as or similar to the other numbered component.

[0108] DETAILED DESCRIPTION OF THE INVENTION

[0109] Purpose

[0110] The methods described herein can be used for many applications and fields of use. For example, high-throughput single-nucleus and single-cell methods can be used for drug discovery. Measurement of transcriptional diversity in single cells induced by drugs or genetic perturbations can be used for drug screening. In one application, genome editing is used to classify genetic and genomic variants (Findlay, Nature, 2018 562(7726):217-222). In another application, the methods can be used to understand health and disease, science, medicine, diagnostics, biomedical research, clinical applications, or biomarker discovery.

[0111] Exposure of cells to predetermined conditions

[0112] The methods provided herein can be used to generate sequencing libraries from a plurality of single cells. In one embodiment, the method includes exposing cells to different predetermined conditions. The method includes exposing subsets of cells to different predetermined conditions (FIG. 1, block 10). The different conditions can include, for example, different culture conditions (e.g., different media, different environmental conditions), different doses of a drug, different drugs, or drug combinations. Drugs are described herein. Nuclei or cells from each subset of cells and / or samples are tagged using nuclear hashing, pooled, and analyzed by a massively multiplexed single-nucleus or single-cell sequencing method. Essentially, single-nucleus or single-cell sequencing methods can be used for, but are not limited to, single-nucleus transcriptome sequencing (U.S. Patent Application No. 62 / 680,259 and Gunderson et al., (WO 2016 / 130704)), single-nucleus whole-genome sequencing (U.S. Patent Application Publication No. 2018 / 0023119), or single-nucleus sequencing of transposon-accessible chromatin (U.S. Patent No. 10,059,989), sciHiC (Ramani et al., Nature Methods, 2017, 14:263-266), DRUG-seq (Ye et al., Nature Commun., 9, article number 4307), or any combination of DNA, RNA, and protein analytes, such as sci-CAR (Cao et al., Science, 2018, 361(6409):1380-1385). Nuclear hashing is used to demultiplex and distinguish individual cells or nuclei from different conditions.

[0113] The cells can be from any organism and can be obtained from any cell type or any tissue of the organism. Typically, the cells are distributed into a first plurality of compartments. In one embodiment, the compartments are wells of a multi-well device, such as a 96-well, 384-well, or 1536-well plate, and the cells within the compartments are referred to as subsets of cells. In one embodiment, the cells within the subsets are genetically homogeneous; in another embodiment, the cells within the subsets are genetically heterogeneous. The number of cells can vary and may depend on the practical limitations of the equipment used in other steps of the methods described herein (e.g., the size of the wells in the multi-well plate, the number of indexes). In one embodiment, the number of cells in the subset may be 100,000,000 or less, 10,000,000 or less, 1,000,000 or less, 100,000 or less, 45,000 or less, 35,000 or less, 25,000 or less, 15,000 or less, 5,000 or less, 1,000 or less, 500 or less, or 50 or less.

[0114] In one embodiment, each subset of cells is exposed to a drug or perturbation. The drug can be essentially anything that causes a change in the cell. For example, the drug can alter the cell's transcriptome, alter the cell's chromatin structure, change the activity of a protein in the cell, modify the cell's DNA, or alter the cell's DNA editing. Examples of drugs include, but are not limited to, compounds such as proteins (including antibodies), non-ribosomal proteins, polyketides, organic molecules (including organic molecules of 900 daltons or less), inorganic molecules, RNA or RNAi molecules, carbohydrates, glycoproteins, nucleic acids, or combinations thereof. In one embodiment, the drug causes a genetic perturbation, e.g., a DNA editing protein such as CRISPR or Talen. In one embodiment, the drug is a drug, such as a therapeutic agent. In one embodiment, the cells can be wild-type cells, or in another embodiment, the cells can be genetically modified to include a genetic perturbation, e.g., a gene knock-in, a gene knock-out, or overexpression (Szlachta et al., Nat Commun., 2018, 9:4275). In one embodiment, cells are modified by introducing genetic perturbations, exposure to drugs, and any resulting changes in, for example, transcription of genome organization can be identified in single cells (Datlinger et al., Nature Methods, 2017, 14(3):297-30 [sci Crop]; Adhemar et al., Cell, 2016, 167:1883-1896 [sci Crispr]; and Dixit et al., Cell, 2016, 167(7):1853-1866) [sci Perturb]. Optionally, in these embodiments using guide RNAs targeting genetic perturbations, the method may further include identifying and confirming the actual edits to the genome.

[0115] Subsets of cells can be exposed to the same drug, but different variables can be varied across the compartments of a multi-well device, allowing multiple variables to be tested in a single experiment. For example, different doses, different exposure periods, and different cell types can be tested in a single plate. In one embodiment, cells can express proteins with known activity, and the effect of drugs on activity can be evaluated under different conditions. Nuclear hashing, used to label nuclei, allows for the subsequent identification of nucleic acids derived from specific subsets of nuclei or cells, for example, from a single well of a multi-well plate.

[0116] Cell and Nuclear Hashing

[0117] In generating a sequencing library from multiple cells or multiple single nuclei, cells or nuclei can be contacted with hashing oligos. The use of hashing oligos is optional and can be used in conjunction with normalization oligos. Normalization oligos are described herein. In one embodiment, contacting occurs after the cells have been exposed to any predetermined conditions. Typically, contacting with hashing oligos occurs when nuclei or cells are separated into multiple compartments. In one embodiment, nuclei can be isolated (FIG. 1, block 11) and labeled with hashing oligos (FIG. 1, block 12). In another embodiment, nuclei are exposed to hashing oligos before separation from the cells. The inventors have determined that any disruption of the cell membrane allows for labeling of nuclei with hashing oligos. Thus, nuclei can be labeled in the presence or absence of cytoplasmic material. Nuclei may be, and typically are, permeabilized during the hashing and / or normalization oligo labeling process; for example, nuclei may be permeabilized before, during, or after labeling of nuclei with hashing oligos and / or normalization oligos. Methods for permeabilizing membranes are known in the art. In another embodiment, the cells are contacted with a hashing oligo.

[0118] Nuclei isolation is achieved by incubating cells in the cell lysis buffer for at least 1 to 20 minutes, such as 5, 10, or 15 minutes. Optionally, cells can be exposed to an external force, such as movement via a pipette, to aid in lysis. An example of a cell lysis buffer includes 10 mM Tris-HCl, pH 7.4, 10 mM NaCl, 3 mM MgCl, 0.1% IGEPAL CA-630, and 1% SUPERase InRNase inhibitor. Those skilled in the art will recognize that the concentrations of these components can be varied to some extent without reducing the usefulness of the cell lysis buffer for isolating nuclei. Those skilled in the art will recognize that RNAse inhibitors, BSA, and / or detergents can be useful in buffers used for nuclei isolation, and other additives can be added to buffers for other downstream single-cell combinatorial indexing applications.

[0119] In one embodiment, the cells or nuclei are processed using any sci-seq method to measure analytes from the cells, including but not limited to DNA, RNA, protein, or a combination thereof.

[0120] In one embodiment, nuclei are isolated from individual cells that are adherent or in suspension. Methods for isolating nuclei from individual cells are known to those skilled in the art. In one embodiment, nuclei are isolated from cells present in a tissue. Methods for obtaining isolated nuclei typically include the steps of preparing the tissue and isolating nuclei from the prepared tissue. In one embodiment, all steps are performed on ice.

[0121] Tissue preparation can involve flash-freezing the tissue in liquid nitrogen and then reducing the size of the tissue to 1 mm or less in diameter by either mincing or applying blunt force. Optionally, cold proteases and / or other enzymes can be used to disrupt intercellular connections. Mincing can be accomplished with a blade to cut the tissue into small pieces. Blunt force can be achieved by crushing the tissue with a hammer or similar object, and the resulting composition of the crushed tissue is called a powder.

[0122] Conventional tissue nuclei extraction techniques typically involve incubating tissues with tissue-specific enzymes (e.g., trypsin) at high temperatures (e.g., 37°C) for 30 minutes to several hours, followed by lysing cells with a cell lysis buffer. The nuclei isolation method described herein and in U.S. Provisional Patent Application No. 62 / 680,259 offers several advantages. (1) No artificial enzymes are introduced, and all steps are performed on ice. This reduces potential perturbations to cellular states (e.g., transcriptome, chromatin, or methylation states). (2) It has been validated across the widest range of tissue types, including disease samples such as brain, lung, kidney, spleen, heart, cerebellum, and tumor tissue. Compared to conventional tissue nuclei extraction techniques that use different enzymes for different tissue types, our new technique potentially reduces bias when comparing cellular states from different tissues. (3) This method also reduces costs and increases efficiency by eliminating the enzyme treatment step. (4) Compared to other nuclear extraction techniques (e.g., Dounce tissue grinders), this technique is more robust to different tissue types (e.g., the Dounce method requires optimizing the Dounce cycle for different tissues) and is capable of processing large sample pieces at high throughput (e.g., the Dounce method is limited by the size of the grinder).

[0123] Optionally, the isolated nuclei may be free of nucleosomes or may be subjected to conditions that deplete the nuclei of nucleosomes, producing nucleosome-depleted nuclei (e.g., U.S. Patent Application Publication No. 2018 / 0023119). Nucleosome-depleted nuclei may be useful for genome-wide single-nucleus sequencing.

[0124] The lysis buffer may contain hashing oligos used for nuclei hashing (Figure 1, block 12). Alternatively, hashing oligos may not be present in the lysis buffer but may be present at a later step deemed appropriate by one of skill in the art.

[0125] Hashing oligos comprise single- or double-stranded nucleic acid sequences, including DNA, RNA, or a combination thereof. Hashing oligos can be DNAse-resistant or RNAse-resistant. Hashing oligos can include any combination of nucleic acid components, including, but not limited to, indexes, UMIs, and universal sequences. Hashing oligos can also include other non-nucleic acid components, including proteins such as antibodies. In one embodiment, a hashing oligo comprises a 5' region, a subset-specific index sequence, and a 3' terminal sequence. The 5' region may be a 5' PCR handle or universal sequence that can be used in subsequent steps for amplifying and adding specific nucleotides to the hashing oligo. The 5' PCR handle may comprise a nucleotide sequence identical to or complementary to a universal capture sequence. The 3' terminal sequence may be any sequence of nucleotides useful in downstream steps. For example, if the downstream step involves generating a transcriptome library, the 3' terminal sequence may comprise a polyadenylation sequence. In another embodiment, the hashing oligos comprise nucleic acid sequences that can be used in a subsequent ligation step, amplification step, primer extension step, or combination thereof to add subset-specific index sequences and other nucleotides, such as a 5' region and / or a polyadenylated 3' end, useful in subsequent steps of the method. Any further manipulation of the hashing oligos to add subset-specific index sequences, and any other elements, can occur before the pooling step described herein. In one embodiment, the hashing oligos can be added before, during, or after the cell lysis step or drug exposure. In one embodiment, the hashing oligos can be added without a cell lysis step.

[0126] The index sequence, also called a tag or barcode, is useful as a marker characteristic of the compartment in which a particular target nucleic acid is present. Thus, the index is a nucleic acid sequence tag attached to each target nucleic acid present in a particular compartment, and the presence of this index indicates or is used to identify the compartment in which the nucleus or cell population resides at this stage of the method. The subset-specific index sequence of the hashing oligo indicates which nucleic acids originate from a subset of cells exposed to a particular predetermined condition, allowing those nucleic acids to be distinguished from nucleic acids originating from other subsets of cells exposed to a particular predetermined condition.

[0127] As used herein, index sequences (e.g., subset-specific index sequences for hashing oligos, population-identifying index sequences for normalization oligos, or other indexes) can be any suitable number of sequences, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more in length. A four-nucleotide tag provides the potential for multiplexing 256 samples, while a six-base tag allows for the processing of 4096 samples. Indexes or barcodes can be introduced through many different methods, including, but not limited to, oligos, ligation, extension, adsorption, and specific or nonspecific interactions of oligos, or amplification.

[0128] The hashing oligos bind to the nuclei or cells of the compartment and are optionally immobilized thereto. Binding can be specific or non-specific. Hashing oligos that specifically bind to cells or nuclei can include any domain that mediates specific binding. Examples of domains include ligands for receptors on the surface of nuclei or cells, antibodies or antibody fragments, aptamers, or specific oligo sequences. The inventors have determined that the hashing oligos bind non-specifically to nuclei in sufficient amounts for subsequent demultiplexing in the method. Thus, in one embodiment, the hashing oligos non-specifically bind to the nuclei or cells of the compartment. In one embodiment, the non-specific binding is due to absorption. In one embodiment, the hashing oligos can be added at any step before pooling. In one embodiment, the hashing oligos are a mixture of one or more oligos, each with a unique barcode or unique sequence. The hashing oligos can include a barcode or a sequence for introducing a barcode at a later stage after binding. In one embodiment, a uniquely barcoded hashing oligo combination can be a signature for a particular experimental condition, agent, sample, or perturbation.

[0129] The unique hashing oligos present in each subset are fixed to the nuclei or cells of that subset by exposure to a cross-linking compound. Useful examples of cross-linking compounds include, but are not limited to, paraformaldehyde, formalin, or methanol. Other useful examples are described in Hermanson (Bioconjugate Techniques, 3rd ed., 2013). Paraformaldehyde can be at a concentration of 1% to 8%, such as 5%. Treatment of nuclei with paraformaldehyde can include adding paraformaldehyde to a suspension of nuclei and incubating at 0°C. In some embodiments, the hashing oligos are not cross-linked but remain bound.

[0130] The manipulation of nuclei or cells, including the pooling and partitioning steps described herein, can include the use of a nuclear buffer. An example of a nuclear buffer is 10 mM Tris-HCl, pH 7.4, 10 mM NaCl, 3 mM MgCl, 1% SUPERase InRNase inhibitor (20 U / μL, Ambion), and 1% BSA (20 mg / mL, NEB). Those skilled in the art will recognize that the concentrations of these components can be varied to some extent without reducing the usefulness of the nuclear buffer for suspending nuclei.

[0131] Isolated, fixed nuclei or cells can be immediately aliquoted and flash-frozen in liquid nitrogen for later use. When prepared for use after freezing, thawed nuclei can be permeabilized, for example, with 0.2% Triton X-100 on ice for 3 minutes and briefly sonicated to reduce nuclear clumping.

[0132] The method further includes pooling subsets of nuclei or cells and subsequently distributing the pooled nuclei or cells into a second plurality of compartments (FIG. 1, block 13). The number of nuclei or cells present in a subset, and therefore in each compartment, can be at least one. The number of nuclei or cells in a subset is not intended to be limiting and could reach billions. In one embodiment, the number present in a subset is 100,000,000 or less, 10,000,000 or less, 1,000,000 or less, 10,000 or less, 4,000 or less, 3,000 or less, 2,000 or less, or 1,000 or less. In one embodiment, the number of nuclei present in a subset can be 1 to 1,000, 1,000 to 10,000, 10,000 to 100,000, or 100,000 to 1,000,000, or 1,000,000 to 10,000,000, or 10,000,000 to 100,000,000. In one embodiment, each compartment can be a well of a multi-well plate, such as a 96- or 384-well plate. In one embodiment, each compartment can be a droplet. Methods for distributing nuclei into subsets are known and routine to those skilled in the art. Fluorescence-activated cell sorting (FACS) cytometry can be used, although simple dilution can also be used. In one embodiment, FACS cytometry is not used.

[0133] The number of compartments in the first dispensing step ( FIG. 1 , block 13) can depend on the format used. For example, the number of compartments can be 2 to 96 compartments (if a 96-well plate is used), or 2 to 384 compartments (if a 384-well plate is used). In one embodiment, multiple plates can be used. For example, at least 2, at least 3, at least 4, etc. compartments can be used from a 96-well plate, or at least 2, at least 3, at least 4, etc. compartments can be used from a 384-well plate. When the type of compartment used is droplets containing two or more nuclei or cells, any number of droplets can be used, such as at least 10,000, at least 100,000, at least 1,000,000, or at least 10,000,000 droplets.

[0134] After nuclei or cells are labeled with hashing oligos, pooled, and partitioned into subsets, different procedures can be used to ultimately generate libraries of distinct nucleic acids within the nuclei or cells and sequence the nucleic acids (Figure 1, block 14). In one embodiment, a single-nuclei transcriptome library is generated as described in detail herein (see Example 1), and in another embodiment, a single-nuclei accessible chromatin library is generated as described in detail herein (see Example 2). However, the procedures used after labeling nuclei with hashing oligos are not intended to be limiting. For example, libraries generated using single-cell combinatorial indexing methods, such as whole-cell single-nuclei libraries, transposon-accessible chromatin libraries, or whole-cell single-nuclei libraries, can be generated using nuclei or cells hashed as described herein.

[0135] Canonical Hashing

[0136] In generating a sequencing library from multiple cells or multiple single nuclei, cells or nuclei can be contacted with a population of normalization oligos. The use of normalization oligos is optional and can be used in conjunction with cell or nuclei hashing, as described herein. Contact with normalization oligos can occur before or after the cells are exposed to any conditions. Contact with the normalization oligo population can occur when the nuclei or cells are in bulk before being separated into compartments. Alternatively, nuclei can be isolated, separated into compartments (Figure 2, block 20), and labeled with a population of normalization oligos (Figure 2, block 22). In another embodiment, nuclei are exposed to normalization oligos before being separated from the cells. As with hashing oligos, any disruption of the cell membrane allows for labeling of nuclei with normalization oligos. Thus, nuclei can be labeled in the presence or absence of cytoplasmic material. Nuclei may be, and typically are, permeabilized during the process of labeling with normalization oligos; for example, nuclei may be permeabilized before, during, or after labeling of nuclei with normalization oligos. Methods for permeabilizing membranes are known in the art. In another embodiment, the cells are contacted with a normalizing oligo.

[0137] Methods such as nuclear isolation, tissue preparation, and nuclear extraction described herein for cell and nuclear hashing can be used for normalized hashing. The lysis buffer can include a population of normalization oligos used for normalized hashing (FIG. 2, block 22). Alternatively, the normalization oligos may not be present in the lysis buffer but may be present at a later step deemed appropriate by those skilled in the art. Those skilled in the art will understand that if both hashing oligos and normalization oligos are used, the cells or nuclei can be exposed to the hashing oligos and normalization oligos at the same time or at different times during the method.

[0138] Normalization hashing serves a different purpose than cell or nucleus hashing. Hashing oligos are typically used to label multiple subsets of cells or nuclei, with cells or nuclei within each subset labeled with a single unique oligo. The identity of the unique hashing oligo associated with a cell or nucleus is captured during sequencing of a library derived from the cells or nuclei, allowing subsequent identification of nucleic acids in the library derived from a specific subset of nuclei or cells, for example, from a single well of a multi-well plate. In contrast, normalization hashing can be used as a standardization method, such as a method for standardizing sequence libraries. In normalization, hashed cells or nuclei are exposed to a composition of normalization oligos, either in bulk or as a subset, which contains multiple normalization oligo populations. The identity of the normalization oligo population associated with a cell or nucleus is captured and counted during sequencing of a library derived from the cells or nuclei, and the counts can be used as an external standard to remove technical noise in cell-to-cell variations in variables such as gene expression. Therefore, normalization oligos can be used as standards to assess the sensitivity and quantitative accuracy of sequencing libraries. This type of standard can evaluate the influence of technical variables, benchmark biological tools, and improve the accurate analysis of samples.

[0139] Normalization oligos typically have the same properties as hashing oligos. Normalization oligos contain single- or double-stranded nucleic acid sequences, including DNA, RNA, or a combination thereof. Normalization oligos can be DNAse-resistant or RNAse-resistant. Normalization oligos can contain any combination of nucleic acid components, such as indexes, UMIs, and universal sequences, but are not limited to these. Normalization oligos can also contain other non-nucleic acid components, including proteins such as antibodies. In one embodiment, normalization oligos contain a 5' region, a population-specific index sequence, and a 3' terminal sequence. The 5' region can be a 5' PCR handle or universal sequence that can be used in subsequent steps for amplifying and adding specific nucleotides to the normalization oligo. The 5' PCR handle can contain a nucleotide sequence identical to or complementary to a universal capture sequence. The 3' terminal sequence can be any series of nucleotides useful in downstream steps. For example, if the downstream step involves generating a transcriptome library, the 3' terminal sequence can contain a polyadenylation sequence. In another embodiment, the normalization oligos comprise nucleic acid sequences that can be used in a subsequent ligation step, amplification step, primer extension step, or combination thereof to add subset-specific index sequences and other nucleotides, such as a 5' region and / or a polyadenylated 3' end, useful in subsequent steps of the method. Any further manipulation of the normalization oligos to add subset-specific index sequences, and any other elements, can be performed before the pooling step described herein. In one embodiment, the normalization oligos can be added before, during, or after the cell lysis step or drug exposure. In one embodiment, the normalization oligos can be added without a cell lysis step.

[0140] The composition of normalization oligos exposed to multiple cells or nuclei comprises multiple distinct populations of normalization oligos. The number of distinct populations can be at least two, and although there is no theoretical upper limit to the number of distinct populations that can be present in the composition, practical considerations such as the cost of producing multiple normalization oligos and the computational time required to analyze many different normalization oligos may limit the number of distinct populations. Without intending to be limiting, the maximum number of populations present in the composition can be 2, 4, 8, 16, 24, 32, 40, 48, 56, 64, 72, 80, 88, 96, or 100. For example, the number of populations present in the composition can be at least 2, at least 4, at least 8, at least 16, at least 24, at least 32, at least 40, at least 48, at least 56, at least 64, at least 72, at least 80, at least 88, or at least 96, and 10 or less, 96 or less, 88 or less, 80 or less, 72 or less, 64 or less, 56 or less, 48 ​​or less, 40 or less, 32 or less, 24 or less, 16 or less, 8 or less, 4 or less, in any combination.

[0141] In one embodiment, each normalization oligo in a single population contains a unique index sequence that is not present in any other population in the composition. In one embodiment, the normalization oligos in a single population are present at one concentration, and the other populations are present at different concentrations, e.g., the concentrations of at least two of the populations are different, and in one embodiment, the concentrations of each population are different. Thus, a composition of normalization oligos can contain multiple populations of oligos, each population containing a different index sequence and each population being the same, or a composition of normalization oligos can contain multiple populations of oligos, each population containing a different index sequence, but at least two populations being different in concentration. In embodiments where each population contains a different index sequence, the concentration of one or more populations is different, but the relationship between each index sequence and its concentration is known.

[0142] In another embodiment, each normalization oligo in a single population contains at least two unique index sequences that are not present in any other population in the composition. For example, one population contains index sequences 1-4, a second population contains index sequences 5-8, and a third population contains index sequences 9-12. Thus, in one embodiment, a normalization oligo composition can contain multiple populations of oligos, each containing a different set of index sequences, with each population having the same concentration. In another embodiment, a normalization oligo composition can contain multiple populations of oligos, each containing a different set of index sequences, but with different concentrations in the other populations, e.g., with different concentrations in at least two of the populations, and in one embodiment, with different concentrations in each population. In embodiments where each population contains a different set of index sequences, the concentrations of one or more populations are different, but the relationship between each index sequence and its concentration is known.

[0143] The concentration of normalization oligos in the composition can vary, depending in part on the capture efficiency of the normalization oligos by cells or nuclei. Generally, capture of normalization oligos is determined by factors including, but not limited to, sample processing and sequencing depth. For mammalian cell lines, the inventors have found that capture rate efficiencies can be low while still yielding useful data (Example 3). It has been empirically determined that normalization oligo compositions can be constructed to capture approximately 6 million normalization oligos per nucleus, resulting in a median UMI count of 1,000-5,000. Those skilled in the art can determine the capture rate efficiency of any cell type and normalization oligo concentration required in the composition to obtain useful normalized data. While not intended to be limiting, the concentration of normalization oligo used is one that results in oligos binding to cells in amounts similar to the amount of cellular analyte (e.g., DNA, RNA, protein, etc.) being measured. In one embodiment, the concentration of each population in the normalization oligo composition can be selected from a concentration of at least 0.001 zeptomole to 100 attomoles or less. For example, the concentration can be at least 0.001 zeptomole, at least 0.01 zeptomole, at least 0.1 zeptomole, at least 1 attomole, or at least 10 attomoles, and up to 100 attomoles, 10 attomoles or less, 1 attomoles or less, 0.1 zeptomole or less, or 0.01 zeptomole or less, in any combination. In one embodiment, the concentration of the population at the lowest level and the concentration of the population at the highest level vary by 1, 2, 3, 4, 5, or 6 orders of magnitude.

[0144] The normalization oligos bind to nuclei or cells and are optionally immobilized to the nuclei or cells. The binding and immobilization of hashing oligos described herein can be used for normalization oligos. For example, binding of normalization oligos can be specific or non-specific. In one embodiment, non-specific binding is by absorption. In one embodiment, normalization oligos that specifically bind to cells or nuclei can include any domain that mediates specific binding. A population of normalization oligos can be immobilized to nuclei or cells as described herein. In some embodiments, the normalization oligos are not cross-linked but remain bound.

[0145] Manipulation of nuclei or cells associated with hashing oligos, normalization oligos, or combinations thereof, including the pooling and partitioning steps described herein, can include the use of a nuclear buffer. An example of a nuclear buffer is 10 mM Tris-HCl, pH 7.4, 10 mM NaCl, 3 mM MgCl, 1% SUPERase InRNase inhibitor (20 U / μL, Ambion), and 1% BSA (20 mg / mL, NEB). Those skilled in the art will recognize that the concentrations of these components can be varied to some extent without reducing the usefulness of the nuclear buffer for suspending nuclei.

[0146] In embodiments in which the normalization oligos are added to cells or nuclei in bulk, the method may further include distributing the pooled cells or nuclei into multiple compartments. Alternatively, in embodiments in which the normalization oligos are added to cells or nuclei present in subsets (FIG. 2, block 22), the method may further include pooling subsets of nuclei or cells and subsequently distributing the pooled nuclei into a second plurality of compartments (FIG. 2, block 24). The number of nuclei or cells present in a subset, and therefore in each compartment, may be at least one. The number of nuclei or cells in a subset is not intended to be limiting and could reach billions. In one embodiment, the number present in a subset is 100,000,000 or less, 10,000,000 or less, 1,000,000 or less, 10,000 or less, 4,000 or less, 3,000 or less, 2,000 or less, or 1,000 or less. In one embodiment, the number of nuclei present in a subset can be 1 to 1,000, 1,000 to 10,000, 10,000 to 100,000, or 100,000 to 1,000,000, or 1,000,000 to 10,000,000, or 10,000,000 to 100,000,000. In one embodiment, each compartment can be a well of a multi-well plate, such as a 96- or 384-well plate. In one embodiment, each compartment can be a droplet. Methods for distributing nuclei into subsets are known and routine to those skilled in the art. Fluorescence-activated cell sorting (FACS) cytometry can be used, although simple dilution can also be used. In one embodiment, FACS cytometry is not used.

[0147] The number of compartments in the first dispensing step ( FIG. 2 , block 24) can depend on the format used. For example, the number of compartments can be 2 to 96 compartments (if a 96-well plate is used), or 2 to 384 compartments (if a 384-well plate is used). In one embodiment, multiple plates can be used. For example, at least 2, at least 3, at least 4, etc. compartments can be used from a 96-well plate, or at least 2, at least 3, at least 4, etc. compartments can be used from a 384-well plate. When the type of compartment used is droplets containing two or more nuclei or cells, any number of droplets can be used, such as at least 10,000, at least 100,000, at least 1,000,000, or at least 10,000,000 droplets.

[0148] After labeling the nuclei or cells with normalization oligos and dividing them into useful or desired subsets, different procedures can be used to ultimately generate a library of different nucleic acids within the nuclei or cells and sequence the nucleic acids (Figure 2, block 26). The procedures used after labeling the nuclei with normalization oligos are not intended to be limiting.

[0149] Single-cell combinatorial indexing of transcriptomes

[0150] The following description of the single-cell combinatorial sequencing method is directed to seq-RNA and is not intended to be limiting. In one embodiment, the method includes indexing mRNA nucleic acids of distributed nuclei (Figure 3, block 30). This step also adds an index to the existing oligos, e.g., hashing oligos, normalization oligos, or both. This index, which is different from the subset-specific index present on the hashing oligos and the population-specific index present on the normalization oligos, is referred to as the first index. Thus, after this step, nucleic acids derived from mRNA molecules have the first index, nucleic acids derived from hashing oligos contain the subset-specific index and the first index, and nucleic acids derived from normalization oligos contain the population-specific index and the first index. In one embodiment, generating nuclei containing the first index involves using reverse transcriptase with an oligo-dT primer to add the index, a random nucleotide sequence, and a universal sequence. The random sequence is used as a unique molecular identifier (UMI) to label unique nucleic acid fragments. The random sequence can also be used to assist in the removal of duplicates in downstream processing. The universal sequence serves as a complementary sequence for hybridization in the ligation step described herein. Exposing nuclei to these components under conditions suitable for reverse transcription produces a population of indexed nuclei, each containing two populations of indexed nucleic acid fragments. One population results from reverse transcription of a hashing or normalization oligo hybridized to an oligo-dT primer, and the other population results from reverse transcription of mRNA nucleic acid hybridized to an oligo-dT primer. The indexed nucleic acid fragments can, and typically do, contain an index sequence on the synthesized strand that indicates a specific section.

[0151] Indexed nuclei from multiple compartments can be combined (FIG. 3, block 31). For example, indexed nuclei from 2 to 24 compartments, 2 to 96 compartments (if a 96-well plate is used), or 2 to 384 compartments (if a 384-well plate is used) are combined. In one embodiment, indexed nuclei from multiple plates are combined. For example, at least 2, at least 3, at least 4, etc. compartments are combined from a 96-well plate, or at least 2, at least 3, at least 4, etc. compartments are combined from a 384-well plate. In one embodiment, four compartments are combined from a 386-well plate. A subset of these combined indexed nuclei, referred to herein as pooled indexed nuclei, is then distributed into a third plurality of compartments (FIG. 3, block 31). The number of nuclei present in a subset, and therefore the number of nuclei in each compartment, is based in part on the desire to reduce index collisions, which are the presence of two nuclei with the same index that end up in the same compartment at this step of the method. In one embodiment, the number of nuclei present in each subset is approximately equal.The method for dividing nuclei into subsets is known and routine for those skilled in the art.Examples include, but are not limited to, simple dilution.In one embodiment, FACS cytometry is not used.

[0152] Following partitioning into subsets of nuclei, the indexed nucleic acid fragments within each compartment are incorporated with a second index sequence to generate doubly indexed fragments, which results in further indexing of the indexed nucleic acid fragments (Figure 3, block 32).

[0153] In one embodiment, incorporating a second index sequence involves ligating a hairpin ligation duplex to the indexed nucleic acid fragment in each section. Using a hairpin ligation duplex to introduce a universal sequence, an index, or a combination thereof, into the end of a target nucleic acid fragment typically involves using one end of the duplex as a primer for subsequent amplification. In contrast, in one embodiment, the hairpin ligation duplex used herein does not act as a primer. An advantage of using the hairpin ligation duplex described herein is the reduction of self-ligation observed with many hairpin ligation duplexes described in the art. In one embodiment, the ligation duplex contains five elements: 1) a universal sequence that is the complement of the universal sequence present on the oligo-dT primer, 2) a second index, 3) ideoxy U, 4) a nucleotide sequence capable of forming a hairpin, and 5) the reverse complement of the second index. The second index sequence is unique to each compartment in which the distributed indexed nuclei are located after the first index is added by reverse transcription (Figure 3, block 31).

[0154] The doubly-indexed nuclei from multiple compartments can be combined (FIG. 3, block 33). For example, doubly-indexed nuclei from 2 to 24 compartments, 2 to 96 compartments (if a 96-well plate is used), or 2 to 384 compartments (if a 384-well plate is used) are combined. In one embodiment, doubly-indexed nuclei from multiple plates are combined. For example, at least 2, at least 3, at least 4, etc. compartments are combined from a 96-well plate, or at least 2, at least 3, at least 4, etc. compartments are combined from a 384-well plate. In one embodiment, 4 compartments are combined from a 386-well plate. A subset of these combined doubly-indexed nuclei, referred to herein as pooled doubly-indexed nuclei, are then distributed into a fourth plurality of compartments (FIG. 3, block 33). The number of nuclei present in a subset, and therefore the number of nuclei in each of its compartments, is based in part on the desire to reduce index collisions, where a collision is the presence of two nuclei with the same transposase index ending up in the same compartment at this step of the method. In one embodiment, 100 to 30,000 nuclei are distributed into each well. In one embodiment, the number of nuclei in a well is at least 100, at least 500, at least 1,000, or at least 5,000. In one embodiment, the number of nuclei in a well is 30,000 or less, 25,000 or less, 20,000 or less, or 15,000 or less. In one embodiment, the number of nuclei present in a subset can be 100 to 1,000, 1,000 to 10,000, 10,000 to 20,000, or 20,000 to 30,000. In one embodiment, 2,500 nuclei are distributed into each well. In one embodiment, the number of nuclei present in each subset is approximately equal.The method for dividing nuclei into subsets is known and routine for those skilled in the art.Examples include, but are not limited to, simple dilution.In one embodiment, FACS cytometry is not used.

[0155] Following partitioning of the doubly indexed nuclei into subsets, the second DNA strand can be synthesized (Figure 3, block 34). Alternatively, synthesis of the second DNA strand can occur prior to partitioning, for example, in bulk.

[0156] The nuclei can then be tagged (Figure 3, block 35). Each compartment containing doubly indexed nuclei contains a transposome complex. The transposome complex can be added to each compartment before, after, or at the same time as the subset of nuclei is added to the compartment. The transposase complex, which is a transposome bound to a transposase recognition site, can insert the transposase recognition site into a target nucleic acid within the nucleus, a process sometimes referred to as "tagging." In some such insertion events, one strand of the transposase recognition site can be transferred to the target nucleic acid. Such a strand is referred to as the "transfer strand." In one embodiment, the transposome complex comprises a dimeric transposase having two subunits and two non-contiguous transposon sequences. In another embodiment, the transposase comprises a dimeric transposase having two subunits and two non-contiguous transposon sequences. In one embodiment, the 5' end of one or both strands of the transposase recognition site can be phosphorylated.

[0157] Some embodiments may involve the use of hyperactive Tn5 transposase and Tn5-type transposase recognition sites (Goryshin and Reznikoff, J. Biol. Chem., 273:7367 (1998)), or MuA transposase and Mu transposase recognition sites containing R1 and R2 end sequences (Mizuuchi, K., Cell, 35:785 (1983); Savilahti, H. et al., EMBO J., 14:4893 (1995)). Tn5 mosaic end (ME) sequences can also be used by those skilled in the art, as optimized.

[0158] Further examples of transposition systems that can be used with certain embodiments of the compositions and methods provided herein include Staphylococcus aureus Tn552 (Colegio et al., J. Bacteriol., 183:2384-8, 2001; Kirby C et al., Mol. Microbiol., 43:173-86, 2002), Ty1 (Devine and Boeke, Nucleic Acids Res., 22:3765-72, 1994, and WO 95 / 23875), transposon Tn7 (Craig, NL, Science. 271:1512, 1996; Craig, NL, Curr. Top. Reviews in Microbiol Immunol., 204:27-48, 1996), Tn / O and IS10 (Kleckner N, et al., Curr Top Microbiol Immunol., 204:49-82, 1996), Mariner transposase (Lampe DJ, et al., EMBO J, 15:5470-9, 1996), Tc1 (Plasterk RH, Curr. Topics Microbiol. Immunol., 204:125-43, 1996), P element (Gloor, GB, Methods Mol. BioBiol., 260:97-114, 2004), Tn3 (Ichikawa and Ohtsubo, J Biol. Chem. 265:18829-32, 1990), bacterial insertion sequences (Ohtsubo and Sekine, Curr. Top. Microbiol. Immunol. 204:1-26, 1996), retroviruses (Brown et al., Proc Natl Acad Sci USA 86:2525-9, 1989), and yeast retrotransposons (Boeke and Corces, Annu Rev Microbiol. 43:403-34, 1989). Other examples include IS5, Tn10, Tn903, IS911, and modified forms of transposase family enzymes (Zhang et al. (2009) PLoS Genet. 5:e1000689. Epub 2009 Oct. 16; Wilson C. et al. (2007) J. Microbiol. Methods 71:332-5).

[0159] Other examples of integrases that can be used with the methods and compositions provided herein include retroviral integrases and integrase recognition sequences for such retroviral integrases, e.g., integrases from HIV-1, HIV-2, SIV, PFV-1, and RSV.

[0160] Transposon sequences useful in the methods and compositions described herein are provided in U.S. Patent Application Publication No. 2012 / 0208705, U.S. Patent Application Publication No. 2012 / 0208724, and WO 2012 / 061832. In some embodiments, the transposon sequence comprises a first transposase recognition site, a second transposase recognition site, and an optional index sequence present between the two transposase recognition sites.

[0161] Some transposome complexes useful herein comprise a transposase having two transposon sequences. In some such embodiments, the two transposon sequences are not linked to each other, in other words, the transposon sequences are not contiguous with each other. Examples of such transposomes are known in the art (see, for example, U.S. Patent Application Publication No. 2010 / 0120098).

[0162] In some embodiments, the transposome complex comprises a transposase sequence nucleic acid that binds to two transposase subunits to form a "looped complex" or "looped transposome." In one example, the transposome comprises a dimeric transposase and a transposon sequence. The looped complex can ensure that the transposon is inserted into the target DNA while maintaining the sequence information of the original target DNA without fragmenting the target DNA. As will be appreciated, the looped structure may insert a desired nucleic acid sequence, such as an index, into the target nucleic acid while maintaining the physical connectivity of the target nucleic acid. In some embodiments, the transposon sequence of the looped transposome complex can include a fragmentation site such that the transposon sequence can be fragmented to create a transposome complex containing two transposon sequences. Such transposon complexes are useful for ensuring that adjacent target DNA fragments into which the transposon is inserted receive a code combination that can be unambiguously assembled at a later stage of the assay.

[0163] The transposome complex may optionally include an index sequence, also referred to as a transposase index. The index sequence is present as part of the transposon sequence. Use of the indexed transposome complex results in a target nucleic acid fragment containing an additional index. In one embodiment, the index sequence may be present on the transferred strand, which is the strand of the transposase recognition site that is transferred to the target nucleic acid.

[0164] Thus, tagging can be used to produce nucleic acid fragments with different types of nucleotide sequences at each end. In one embodiment, the resulting nucleic acid fragments contain different nucleotide sequences at each end, such as an N5 primer sequence at one end and an N7 primer sequence at the other, or different universal sequences at each end. Examples of useful universal sequences include, for example, hairpin ligation duplexes and universal sequences to which a universal primer can bind. In other embodiments, tagging can be used to produce nucleic acid fragments with the same type of nucleotide sequence at each end. In one embodiment, the resulting nucleic acid fragments contain nucleotide sequences at each end that have a universal sequence, an index, or both a universal sequence and an index to which a universal primer can bind. The universal sequence can serve as a complementary sequence for hybridization in the amplification step described herein to introduce a third index.

[0165] Following nuclei tagging and nucleic acid fragment processing, a wash process can be performed to enhance molecular purity. Any suitable wash process, such as electrophoresis or size exclusion chromatography, may be used. In some embodiments, solid-phase reversibly immobilized paramagnetic beads can be used to separate desired DNA molecules from unincorporated primers and select nucleic acids based on size, for example. Solid-phase reversibly immobilized paramagnetic beads are commercially available from Beckman Coulter (Agencourt AMPure XP), Thermo Fisher (MagJet), Omega Biotech (Mag-Bind), Promega Beads, and Kapa Biosystems (Kapa Pure Beads).

[0166] Removal of ideoxy U, optionally present in the hairpin region of the hairpin ligation duplex incorporated into the nucleic acid fragment, can occur before, during, or after washing. Removal of uracil residues can be achieved by any available method, and in one embodiment, Uracil-Specific Removal Reagent (USER), available from NEB, is used.

[0167] Following nuclei tagging, a third index sequence can be incorporated into the doubly-indexed nucleic acid fragments in each compartment to generate triply-indexed fragments, where the third index sequence in each compartment is different from the first and second index sequences within the compartment. This results in further indexing of the indexed nucleic acid fragments (Figure 3, block 36) prior to immobilization and sequencing. The third index can be incorporated by an amplification step such as PCR. In one embodiment, universal sequences present at the ends of the doubly-indexed nucleic acid fragments (e.g., a doubly inserted nucleotide sequence at one end and a transposome complex inserted nucleotide sequence at the other end) can be used for primer binding and extension in an amplification reaction. Typically, two different primers are used. One primer hybridizes to the universal sequence at the 3' end of one strand of the doubly-indexed nucleic acid fragment, and the second primer hybridizes to the universal sequence at the 3' end of the other strand of the doubly-indexed nucleic acid fragment. Thus, the anchor sequence present on each primer (e.g., the site to which a universal primer, such as a sequencing primer for Read 1 or Read 2, anneals for sequencing) can be different. Suitable primers can each include an additional universal sequence, such as a universal capture sequence (e.g., the site to which a capture oligonucleotide hybridizes, where the capture oligonucleotide can be immobilized on the surface of a solid substrate). Because each primer contains an index, this step adds another index sequence, one at each end of the nucleic acid fragment, generating triply indexed fragments. In one embodiment, a third index can be added using indexed primers, such as an indexed P5 primer and an indexed P7 primer. The triply indexed fragments can be pooled and subjected to a wash step as described herein.

[0168] The resulting triply indexed fragments collectively provide a library of nucleic acids that can be immobilized and sequenced. The term library, also referred to herein as a sequencing library, refers to the recovery of nucleic acid fragments from a single nucleus that contain known universal sequences at their 3' and 5' ends. In this embodiment, the library contains whole transcriptome nucleic acids from one or more isolated nuclei and can be used to perform whole transcriptome sequencing.

[0169] Preparation of fixed samples for sequencing

[0170] A plurality of multiply indexed fragments can be prepared for sequencing. For example, in an embodiment in which a transcript library of triply indexed fragments is generated, the triply indexed fragments are pooled and typically enriched prior to sequencing by immobilization and / or amplification (Figure 3, block 37). Methods for attaching indexed fragments from one or more sources to a substrate are known in the art. In one embodiment, the indexed fragments are enriched using a plurality of capture oligonucleotides with specificity for the indexed fragments, and the capture oligonucleotides may be immobilized on the surface of a solid substrate. For example, the capture oligonucleotides may comprise a first member of a universal binding pair, and a second member of the binding pair is immobilized on the surface of the solid substrate. Similarly, methods for amplifying immobilized triply indexed fragments include, but are not limited to, bridge amplification and binding equilibrium exclusion. Methods for immobilization and amplification prior to sequencing are described, for example, in Bignell et al. (U.S. Patent No. 8,053,192), Gunderson et al. (WO 2016 / 130704), Shen et al. (U.S. Patent No. 8,895,249), and Pipenburg et al. (U.S. Patent No. 9,309,502).

[0171] Pooled samples can be immobilized in preparation for sequencing. Sequencing can be performed as a single molecule array or can be amplified prior to sequencing. Amplification can be performed using one or more immobilized primers. The immobilized primers can be, for example, on a flat surface or in a lawn on a pool of beads. The pool of beads can be isolated in an emulsion with a single bead in each "compartment" of the emulsion. At a concentration of only one template per "compartment," only a single template is amplified on each bead.

[0172] As used herein, the term "solid-phase amplification" refers to any nucleic acid amplification reaction performed on or in association with a solid support such that all or a portion of the amplification product is immobilized on the solid support upon formation. Specifically, the term encompasses solid-phase polymerase chain reaction (solid-phase PCR) and solid-phase isothermal amplification, which are reactions similar to standard solution-phase amplification except that one or both of the forward and reverse amplification primers are immobilized on the solid support. Solid-phase PCR covers systems such as emulsions in which one primer is immobilized on a bead and the other is in free solution, and colonization of solid-phase gel matrices in which one primer is immobilized on a surface and the other is in free solution.

[0173] In some embodiments, the solid support comprises a patterned surface. A "patterned surface" refers to the arrangement of distinct regions within or on an exposed layer of a solid support. For example, one or more regions can be features in which one or more amplification primers are present. The features can be separated by interstitial regions in which no amplification primers are present. In some embodiments, the pattern can be an xy format of features in rows and columns. In some embodiments, the pattern can be a repetitive sequence of features and / or interstitial regions. In some embodiments, the pattern can be a random sequence of features and / or interstitial regions. Exemplary patterned surfaces that can be used in the methods and compositions described herein are described in U.S. Pat. Nos. 8,778,848, 8,778,849, and 9,079,148, and U.S. Patent Application Publication No. 2014 / 0243224.

[0174] In some embodiments, the solid support comprises an array of wells or depressions on its surface, which can be fabricated as commonly known in the art using a variety of techniques, including but not limited to photolithography, stamping techniques, molding techniques, and microetching techniques. As understood in the art, the technique used will depend on the composition and shape of the array substrate.

[0175] The features within the patterned surface can be wells of an array of wells (e.g., microwells or nanowells) on glass, silicon, plastic, or other suitable solid support with a patterned covalently attached gel, such as poly(N-(5-azidoacetamylpentyl)acrylamide-co-acrylamide) (PAZAM, see, e.g., U.S. Patent Application Publication Nos. 2013 / 184796, WO 2016 / 066586, and 2015 / 002813, each of which is incorporated by reference in its entirety). This process creates a gel pad used for sequencing, which can be stable over many cycles of sequencing operations. Covalently attaching a polymer to the wells is useful for maintaining the gel in the structured features throughout the life of the structured substrate during various applications. However, in many embodiments, the gel need not be covalently attached to the wells. For example, in some conditions, silane-free acrylamide (SFA, see, e.g., U.S. Pat. No. 8,563,477) that is not covalently bonded to any part of the structured matrix can be used as the gel material.

[0176] In certain other embodiments, a structured substrate can be created by patterning a solid support material with wells (e.g., microwells or nanocells), coating the patterned support with a gel material (e.g., PAZAM, SFA, or a chemically modified variant thereof), such as an azido-SFA version, and polishing the gel-coated support, e.g., by chemical or mechanical polishing, thereby retaining the gel within the wells but removing or inactivating substantially all of the gel from the interstitial regions on the surface of the structured substrate between the wells. Primer nucleic acids can be attached to the gel material. A solution of triple-indexed fragments can then be contacted with the polishing substrate, such that individual triple-indexed fragments are seeded into individual wells through interaction with primers attached to the gel material, while target nucleic acids do not occupy the interstitial regions due to the absence or inactivity of the gel material. Amplification of the triple-indexed fragments will be confined to the wells because the absence or inactivity of gel in the interstitial regions prevents outward migration of growing nucleic acid colonies. The process is conveniently manufacturable and scalable, utilizing conventional micro- or nano-fabrication methods.

[0177] Although the present disclosure encompasses "solid-phase" amplification methods in which only one amplification primer is immobilized (the other primer is typically in free solution), in one embodiment, the solid support provides both immobilized forward and reverse primers. In practice, because the amplification process requires an excess of primers to maintain amplification, there will be "multiple" identical forward primers and / or "multiple" identical reverse primers immobilized on the solid support. References herein to forward and reverse primers should be construed as encompassing "multiple" such primers, unless the context dictates otherwise.

[0178] As will be understood by those skilled in the art, any given amplification reaction requires at least one type of forward primer and at least one type of reverse primer specific to the template to be amplified. However, in certain embodiments, the forward and reverse primers may contain template-specific portions of the same sequence and may have the exact same nucleotide sequence and structure (including any non-nucleotide modifications). In other words, solid-phase amplification can be performed using only one type of primer, and such single-primer methods are encompassed within the scope of the present disclosure. Other embodiments may use forward and reverse primers that contain the same template-specific sequence but differ in some other structural features. For example, one type of primer may contain a non-nucleotide modification that is not present in the other.

[0179] Primers for solid-phase amplification are preferably immobilized to the solid support by a single-point covalent bond at or near the 5' end of the primer, leaving the template-specific portion of the primer free to anneal to its cognate template and the 3' hydroxyl group free of primer extension. Any suitable covalent bonding means known in the art can be used for this purpose. The attachment chemistry chosen will depend on the nature of the solid support and any derivatization or functionalization applied to it. The primer itself may contain a moiety, which may be a non-nucleotide chemical modification, to facilitate attachment. In certain embodiments, the primer may contain a sulfur-containing nucleophile, such as a phosphorothioate or thiophosphate, at the 5' end. In the case of solid-supported polyacrylamide hydrogels, this nucleophile binds to bromoacetamide groups present in the hydrogel. A more specific means of attaching primers and templates to the solid support is via a 5' phosphorothioate bond to a hydrogel composed of polymerized acrylamide and N-(5-bromoacetamidoylpentyl)acrylamide (BRAPA), as described in WO 05 / 065814.

[0180] Certain embodiments of the present disclosure may utilize solid supports comprising an inert substrate or matrix (e.g., glass slides, polymeric beads, etc.) that have been "functionalized" by the application of a layer or coating of an intermediate material containing reactive groups that allow for covalent attachment to biomolecules, such as polynucleotides. Examples of such supports include, but are not limited to, polyacrylamide hydrogels supported on an inert substrate such as glass. In such embodiments, the biomolecule (e.g., polynucleotide) may be covalently attached directly to the intermediate material (e.g., hydrogel), or the intermediate material may itself be noncovalently attached to the substrate or matrix (e.g., glass substrate). The term "covalent attachment to a solid support" should be interpreted accordingly to encompass this type of arrangement.

[0181] The pooled samples may be amplified on beads, each bead containing a forward and reverse amplification primer. In certain embodiments, a library of triple-indexed fragments is used to prepare clustered arrays of nucleic acid colonies by solid-phase amplification, more specifically solid-phase isothermal amplification, similar to those described in U.S. Patent Application Publication No. 2005 / 0100900, U.S. Patent No. 7,115,400, WO 00 / 18957, and WO 98 / 44151. The terms "cluster" and "colony" are used interchangeably herein and refer to distinct sites on a solid support containing a plurality of identical immobilized nucleic acid strands and a plurality of identical immobilized complementary nucleic acid strands. The term "clustered array" refers to an array formed from such clusters or colonies. In this context, the term "array" should not be understood as requiring an ordered arrangement of the clusters.

[0182] The terms "solid phase" or "surface" are used to refer to either a planar array in which the primers are attached to a flat surface, such as a glass, silica, or plastic microscope slide, or similar flow cell device, or beads, to which one or two primers are attached and the beads are amplified, or an array of beads on a surface after the beads have been amplified.

[0183] Clustered sequences can be prepared using a process of thermal cycling, such as that described in WO 98 / 44151, or a process in which the temperature is held constant and cycles of extension and denaturation are performed using changes in reagents. Such isothermal amplification methods are described in WO 02 / 46456 and U.S. Patent Application Publication No. 2008 / 0009420. Due to the lower temperatures available in isothermal processes, this is particularly preferred in some embodiments.

[0184] It will be understood that any of the amplification methods described herein or generally known in the art can be used with universal or target-specific primers to amplify immobilized DNA fragments. Suitable methods for amplification include, but are not limited to, polymerase chain reaction (PCR), strand displacement amplification (SDA), transcription-mediated amplification (TMA), and nucleic acid sequence-based amplification (NASBA), as described in U.S. Pat. No. 8,003,354. The above amplification methods can be used to amplify one or more nucleic acids of interest. For example, PCR, including multiplex PCR, SDA, TMA, NASBA, etc., can be used to amplify immobilized DNA fragments. In some embodiments, primers specifically directed to the polynucleotide of interest are included in the amplification reaction.

[0185] Other methods suitable for amplifying polynucleotides may include oligonucleotide extension and ligation, rolling circle amplification (RCA) (Lizardi et al., Nat. Genet. 19:225-232 (1998)), and oligonucleotide ligation assay (OLA) techniques (see generally U.S. Pat. Nos. 7,582,420, 5,185,243, 5,679,524, and 5,573,907; European Patent Nos. 0 320 308 B1, 0 336 731 B1, and 0 336 731 B1). (See, for example, WO 90 / 01069, WO 89 / 12696, and WO 89 / 09835.) It will be understood that these amplification methods can be designed to amplify immobilized DNA fragments. For example, in some embodiments, amplification methods can include ligation probe amplification or oligonucleotide ligation assay (OLA) reactions containing primers specifically directed to the nucleic acid of interest. In some embodiments, amplification methods can include primer extension ligation reactions containing primers specifically directed to the nucleic acid of interest. Non-limiting examples of primer extension and ligation primers that can be specifically designed to amplify the nucleic acid of interest include the primers used in the GoldenGate assay (Illumina, San Diego, CA), as exemplified by U.S. Pat. Nos. 7,582,420 and 7,611,869.

[0186] DNA nanoblocks can also be used in combination with the methods and compositions described herein.Methods for making and using DNA nanoblocks for genome sequencing can be found in, for example, U.S. Patents and Publications U.S. Patent No. 7,910,354, U.S. Patent No. 2009 / 0264299, U.S. Patent No. 2009 / 0011943, U.S. Patent No. 2009 / 0005252, U.S. Patent No. 2009 / 0155781, U.S. Patent No. 2009 / 0118488, and can be found, for example, as described in Drmanac et al. (2010, Science 327(5961):78-81). Briefly, after fragmenting genomic library DNA, adapters are ligated to the fragments, and the adapter-ligated fragments are circularized by ligation with circle ligase, followed by rolling circle amplification (as described in Lizardi et al., 1998, Nat. Genet. 19:225-232 and U.S. Patent Application Publication No. 2007 / 0099208(A1)). The extended concatemeric structure of the amplicon promotes coiling, thereby generating compact DNA nanoballs. DNA nanoballs can be captured on a substrate to form ordered or patterned arrays, preferably such that the distance between each nanoball is maintained, thereby enabling sequencing of individual DNA nanoballs. In some embodiments, such as those used by Complete Genomics (Mountain View, CA), successive rounds of adapter ligation, amplification, and digestion are performed prior to circularization to generate head-to-tail constructs with several genomic DNA fragments separated by adapter sequences.

[0187] Exemplary isothermal amplification methods that can be used in the methods of the present disclosure include, but are not limited to, multiple displacement amplification (MDA), exemplified by, for example, Dean et al., Proc. Natl. Acad. Sci. USA 99:5261-66 (2002), or isothermal strand displacement nucleic acid amplification, exemplified, for example, by U.S. Patent No. 6,214,587. Other non-PCR-based methods that can be used in the present disclosure include, for example, those described in Walker et al., Molecular Examples of suitable isothermal amplification methods include strand displacement amplification (SDA), as described in Methods for Virus Detection, Academic Press, 1995; U.S. Patent Nos. 5,455,166 and 5,130,238; and Walker et al., Nucl. Acids Res. 20:1691-96 (1992); or hyperbranched strand displacement amplification, as described, for example, in Lage et al., Genome Res. 13:294-307 (2003). Isothermal amplification methods can be used, for example, with strand-displacing Phi 29 polymerase or Bst DNA polymerase large fragment, 5'->3' exo for random-primed amplification of genomic DNA. The use of these polymerases takes advantage of their high processivity and strand displacement activity. Due to their high processivity, the polymerases can produce fragments of 10-20 kb in length. As noted above, polymerases with low processivity and polymerases with strand displacement activity, such as Klenow polymerase, can be used to produce smaller fragments under isothermal conditions. Further description of amplification reactions, conditions, and components is provided in detail in the disclosure of U.S. Patent No. 7,670,810.

[0188] Another polynucleotide amplification method useful in the present disclosure is tagged PCR, which uses a population of two-domain primers with a 5' region followed by a random 3' region, as described, for example, in Grothues et al., Nucleic Acids Res. 21(5):1321-2, (1993). The first round of amplification is performed on heat-denatured DNA to allow multiple initiation based on individual hybridization from the randomly synthesized 3' region. Due to the nature of the 3' region, initiation sites are considered random throughout the genome. Unbound primers can then be removed, and further replication can be performed using primers complementary to the constant 5' region.

[0189] In some embodiments, isothermal amplification can be performed using binding equilibrium exclusion amplification (KEA), also known as exclusion amplification (ExAmp). The nucleic acid libraries of the present disclosure can be created using a method comprising reacting amplification reagents to produce a plurality of amplification sites, each containing a substantially clonal population of amplicons from individual target nucleic acids seeded at the sites. In some embodiments, the amplification reaction proceeds until a sufficient number of amplicons are produced to fill each amplification site to capacity. In this manner, filling a previously seeded site to capacity inhibits target nucleic acids from landing and amplifying at that site, thereby producing a clonal population of amplicons at that site. In some embodiments, apparent clonality can be achieved even if the amplification site is not filled to capacity before a second target nucleic acid arrives at the site. Under some conditions, amplification of a first target nucleic acid can proceed to a point where a sufficient number of copies are produced to effectively exceed or overwhelm the production of copies from a second target nucleic acid transported to that site. For example, in an embodiment using a bridge amplification process on circular features less than 500 nm in diameter, it was determined that after 14 cycles of exponential amplification for a first target nucleic acid, contamination from a second target nucleic acid at the same site produces an insufficient number of contaminating amplicons to adversely affect sequence synthesis analysis on an Illumina sequencing platform.

[0190] In some embodiments, amplification sites in an array can be, but are not necessarily, completely clonal. Rather, in some applications, individual amplification sites can be populated primarily with amplicons from a first triplicate-indexed fragment and can also have low levels of contaminating amplicons from a second target nucleic acid. An array can have one or more amplification sites with low levels of contaminating amplicons, as long as the level of contamination does not have an unacceptable effect on subsequent use of the array. For example, if an array is used in a detection application, an acceptable level of contamination is one that does not unacceptably affect the signal-to-noise ratio or resolution of the detection technique. Thus, apparent clonality generally relates to the particular use or application of an array produced by the methods described herein. Exemplary levels of contamination that are acceptable in individual amplification sites for a particular application include, but are not limited to, up to 0.1%, 0.5%, 1%, 5%, 10%, or 25% contaminating amplicons. An array can include one or more amplification sites with these exemplary levels of contaminating amplicons. For example, up to 5%, 10%, 25%, 50%, 75%, or even 100% of the amplification sites within an array may contain contaminating amplicons. It will be understood that in an array or other collection of sites, at least 50%, 75%, 80%, 85%, 90%, 95%, or 99% or more of the sites may be clonal or appear clonal.

[0191] In some embodiments, binding equilibrium exclusion can occur when a process occurs at a rate fast enough to effectively exclude another event or process from occurring. Take the example of creating a nucleic acid array in which array sites are randomly seeded with triply indexed fragments from solution, and copies of the triply indexed fragments are produced in an amplification process to fill each seeded site to capacity. According to the binding equilibrium exclusion method of the present disclosure, the seeding and amplification processes can proceed simultaneously under conditions in which the amplification rate exceeds the seeding rate. Thus, the relatively fast rate at which copies are made at a site seeded by a first target nucleic acid effectively excludes a second nucleic acid from seeding that site for amplification. The binding equilibrium exclusion method can be performed as described in detail in the disclosure of U.S. Patent Application Publication No. 2013 / 0338042.

[0192] Binding equilibrium exclusion can take advantage of a relatively slow rate for initiating amplification (e.g., a slow rate for making the first copy of a triply-indexed fragment) versus a relatively fast rate for making subsequent copies of a triply-indexed fragment (or first copy of a triply-indexed fragment). In the example of the previous paragraph, binding equilibrium exclusion occurs due to a relatively slow rate of triply-indexed fragment seeding (e.g., relatively slow diffusion or transport) versus a relatively fast rate at which amplification occurs to fill the site with triply-indexed fragment species copies. In another exemplary embodiment, binding equilibrium exclusion can occur due to a delay (e.g., delayed or slow activation) in the formation of the first copy of a triply-indexed fragment that seeded the site versus a relatively fast rate at which subsequent copies are made to fill the site. In this example, an individual site may be seeded with several different triply-indexed fragments (e.g., several triply-indexed fragments may be present at each site prior to amplification). However, because first copy formation of any given triply-indexed fragment can be activated randomly, the average rate of first copy formation will be relatively slow compared to the rate at which subsequent copies are generated. In this case, an individual site may be seeded with several different triply indexed fragments, but due to binding equilibrium exclusion, only one of those triply indexed fragments can be amplified. More specifically, when a first triply indexed fragment is activated for amplification, the site is rapidly filled to capacity with copies of that fragment, thereby preventing copies of a second triply indexed fragment from being made at the site.

[0193] In one embodiment, the method is performed simultaneously to (i) triple-index fragments into the amplification site at an average transport rate and (ii) amplify the triple-index fragments at the amplification site at an average amplification rate, where the average amplification rate exceeds the average transport rate (U.S. Patent No. 9,169,513). Therefore, in such an embodiment, binding equilibrium exclusion can be achieved by using a relatively slow transport rate. For example, a sufficiently low concentration of triple-index fragments can be selected to achieve the desired average transport rate, since lower concentrations result in slower transport rates. Alternatively or additionally, a high viscosity solution and / or the presence of a molecular crowding agent in the solution can be used to reduce the transport rate. Examples of useful molecular crowding agents include, but are not limited to, polyethylene glycol (PEG), Ficoll, dextran, or polyvinyl alcohol. Exemplary molecular crowding agents and formulations are described in U.S. Patent No. 7,399,590, incorporated herein by reference. Another factor that can be adjusted to achieve the desired transport rate is the average size of the target nucleic acid.

[0194] Amplification reagents can include additional components that promote amplicon formation, potentially increasing the rate of amplicon formation. One example is a recombinase. Recombinases can promote amplicon formation by enabling repeated infiltration / extension. More specifically, recombinases can promote the infiltration of triple-indexed fragments by a polymerase and the extension of primers by the polymerase using the triple-indexed fragments as templates for amplicon formation. This process can be repeated as a chain reaction, with amplicons produced from each round of infiltration / extension serving as templates in subsequent rounds. Because no denaturation cycles (e.g., by heating or chemical denaturation) are required, this process can be performed more quickly than standard PCR. Therefore, recombinase-promoted amplification can be performed isothermally. It is desirable to include ATP or other nucleotides (or optionally non-hydrolyzable analogs thereof) in the recombinase-promoted amplification reagent to promote amplification. A mixture of recombinase and single-stranded binding (SSB) protein is particularly useful, as SSB can further promote amplification. Representative formulations for recombinase-facilitated amplification include those sold commercially as TwistAmp kits by TwistDx Ltd. (Cambridge, UK). Useful components and reaction conditions for recombinase-facilitated amplification reagents are described in U.S. Patent Nos. 5,223,414 and 7,399,590.

[0195] Another example of a component that can be included in an amplification reagent to promote amplicon formation and, in some cases, increase the rate of amplicon formation is helicase. Helicase can promote amplicon formation by enabling a chain reaction of amplicon formation. Because no denaturation cycles (e.g., by heating or chemical denaturation) are required, this process can be performed more rapidly than standard PCR. Therefore, helicase-promoted amplification can be performed isothermally. A mixture of helicase and single-strand binding (SSB) protein is particularly useful, as SSB can further promote amplification. Representative formulations for helicase-promoted amplification include those commercially available as IsoAmp kits from Biohelix, Inc. (Beverly, Massachusetts). Further examples of useful formulations containing helicase proteins are described in U.S. Patent Nos. 7,399,590 and 7,829,284.

[0196] Yet another example of a component that can be included in an amplification reagent to facilitate amplicon formation, and in some cases increase the rate of amplicon formation, is an origin binding protein.

[0197] Sequencing Uses / Sequencing Methods

[0198] Following attachment of the triply indexed fragments to the surface, the immobilized and amplified triply indexed fragments are sequenced. Sequencing can be performed using any suitable sequencing technique, and methods for determining the sequence of the immobilized and amplified triply indexed fragments, including strand resynthesis, are known in the art and are described, for example, in Bignell et al. (U.S. Patent No. 8,053,192), Gunderson et al. (WO 2016 / 130704), Shen et al. (U.S. Patent No. 8,895,249), and Pipenburg et al. (U.S. Patent No. 9,309,502).

[0199] The methods described herein can be used in conjunction with various nucleic acid sequencing methods. Particularly applicable techniques are those in which nucleic acids are attached to fixed positions within an array so that their relative positions do not change, and the array is repeatedly imaged. For example, embodiments in which images are obtained in different color channels corresponding to different labels used to distinguish one nucleotide base type from another are particularly applicable. In some embodiments, the process of determining the nucleotide sequence of the triple index fragments can be an automated process. Preferred embodiments include sequencing-by-synthesis ("SBS") techniques.

[0200] SBS technology generally involves the enzymatic extension of nascent nucleic acid chain by repeatedly adding nucleotide to template chain.In the conventional method of SBS, a single nucleotide monomer can be provided to target nucleic acid in the presence of polymerase in each delivery.However, in the method described herein, multiple types of nucleotide monomers can be provided to target nucleic acid in the presence of polymerase during delivery.

[0201] In one embodiment, the nucleotide monomer comprises a locked nucleic acid (LNA) or a bridged nucleic acid (BNA). The use of an LNA or BNA in the nucleotide monomer increases the hybridization strength between the nucleotide monomer and the sequencing primer sequence present on the immobilized triplex index fragment.

[0202] SBS can use nucleotide monomers with or without terminator moieties. Methods using nucleotide monomers without terminators include, for example, pyrosequencing and sequencing using γ-phosphate-labeled nucleotides, as described in more detail herein. In methods using nucleotide monomers without terminators, the number of nucleotides added in each cycle is generally variable and depends on the template sequence and the mode of nucleotide delivery. In SBS techniques using nucleotide monomers with terminator moieties, the terminators can be effectively irreversible under the sequencing conditions used, as in traditional Sanger sequencing using dideoxynucleotides, or they can be reversible, as in the sequencing method developed by Solexa (now Illumina).

[0203] SBS techniques can use nucleotide monomers that have a label moiety or lack a label moiety. Therefore, incorporation events can be detected based on the properties of the label, such as the fluorescence of the label, the properties of the nucleotide monomer, such as molecular weight or charge, or by-products of nucleotide incorporation, such as the release of pyrophosphate. In embodiments in which two or more different nucleotides are present in the sequencing reagent, the different nucleotides may be distinguishable from one another, or the two or more different labels may be distinguishable under the detection technique used. For example, different nucleotides present in the sequencing reagent can have different labels, which can be distinguished using appropriate optical systems, as exemplified by the sequencing method developed by Solexa (now Illumina).

[0204] A preferred embodiment includes pyrosequencing technology. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) when a specific nucleotide is incorporated into the nascent strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M., and Nyren, P. (1996) "Real-time DNA sequencing using inorganic pyrophosphate detection", Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001) "Pyrosequencing sheds light on DNA sequencing", Genome Res., 11(1), 3-11; Ronaghi, M., Uhlen, M., and Nyren, P. (1998) "Real-time inorganic pyrophosphate-based sequencing", Science 281 (5375), 363, U.S. Patent Nos. 6,210,891, 6,258,568, and 6,274,320). In pyrosequencing, released PPi can be detected by its immediate conversion to adenosine triphosphate (ATP) by ATP sulfurase, and the level of generated ATP is detected via luciferase-generated photons. The nucleic acids to be sequenced can be attached to features in an array, and the array can be imaged to capture the chemiluminescent signal produced by the incorporation of nucleotides into the features of the array. Images can be obtained after treating the array with a specific nucleotide type (e.g., T, C, or G). The images obtained after the addition of each nucleotide type differ in terms of which features in the array are detected. These differences in the images reflect the different sequence content of the features on the array. However, the relative positions of each feature remain unchanged in the image. Images can be stored, processed, and analyzed using the methods described herein. For example, images obtained after treating the array with each different nucleotide type can be processed in the same manner as exemplified herein for images obtained from different detection channels for reversible terminator-based sequencing methods.

[0205] In another exemplary type of SBS, cycle sequencing is achieved by the stepwise addition of reversible terminator nucleotides containing cleavable or photobleachable dye labels, as described, for example, in WO 04 / 018497 and U.S. Patent No. 7,057,026. This approach has been commercialized by Solexa (now Illumina) and is also described in WO 91 / 06678 and WO 07 / 123,744. The availability of fluorescently labeled terminators, both of whose termini can be reversed and from which the fluorescent labels are cleaved, facilitates efficient cyclic reversible termination (CRT) sequencing. Polymerases can also be co-engineered to efficiently incorporate and extend from these modified nucleotides.

[0206] In some reversible terminator-based sequencing embodiments, the label does not substantially inhibit extension under SBS reaction conditions. However, the detection label may be removable, for example, by cleavage or degradation. Images can be captured after incorporation of the label into arrayed nucleic acid features. In certain embodiments, each cycle involves simultaneous delivery of four different nucleotide types to the array, with each nucleotide type bearing a spectrally distinct label. Four images can then be obtained, each using a detection channel selective for one of the four different labels. Alternatively, different nucleotide types can be added sequentially, with images of the array being obtained between each addition step. In such embodiments, each image shows nucleic acid features incorporating a particular type of nucleotide. Because the sequence content of each feature varies, different features are present or absent in different images. However, the relative positions of the features remain unchanged within the image. Images obtained from such reversible terminator-SBS methods can be stored, processed, and analyzed as described herein. Following the image capture step, the label can be removed, and the reversible terminator moiety can be removed for subsequent cycles of nucleotide addition and detection. Removal of the label after detection in a particular cycle and before the subsequent cycle has the advantage of reducing background signal and crosstalk between cycles. Examples of useful labeling and removal methods are described herein.

[0207] In certain embodiments, some or all of the nucleotide monomers may contain reversible terminators. In such embodiments, the reversible terminator / cleavable fluorophore may comprise a fluorophore attached to the ribose moiety via a 3' ester bond (Metzker, Genome Res. 15:1767-1776 (2005)). Other approaches have separated the terminator chemistry from the fluorescent label (Ruparel et al., Proc Natl Acad Sci USA 102:5932-7 (2005)). Ruparel et al. describe the development of a reversible terminator that uses a small 3' allyl group to block extension but can be easily unblocked by brief treatment with a palladium catalyst. The fluorophore was attached to the group via a photocleavable linker that can be easily cleaved by 30 seconds of exposure to long-wavelength UV light. Thus, either disulfide reduction or photocleavage can be used as the cleavable linker. Another approach to reversible termination is the use of a natural terminator following the placement of a bulky dye on the dNTP. The presence of a charged bulky dye on the dNTP can act as an effective terminator through steric and / or electrostatic hindrance. The presence of one incorporation event prevents further binding unless the dye is removed. Cleavage of the dye removes the fluorophore, effectively reversing the terminus. Examples of modified nucleotides are also described in U.S. Patent Nos. 7,427,673 and 7,057,026.

[0208] Additional exemplary SBS systems and methods that can be utilized with the methods and systems described herein are described in U.S. Patent Application Publication Nos. 2007 / 0166705, 2006 / 0188901, 2006 / 0240439, 2006 / 0281109, 2012 / 0270305, and 2013 / 0260372, U.S. Patent No. 7,057,026, and WO 05 / 065814, U.S. Patent Application Publication No. 2005 / 0100900, and WO 06 / 064199 and WO 07 / 010,251.

[0209] Some embodiments may employ detection of four different nucleotides using fewer than four different labels. For example, SBS may be performed using the methods and systems described in the incorporated document, U.S. Patent Publication No. 2013 / 0079232. As a first example, pairs of nucleotide types may be detected at the same wavelength but may be distinguished based on differences in intensity for one member of the pair or based on a change to one member of the pair (e.g., via chemical, photochemical, or physical modification) that results in the appearance or disappearance of a distinct signal compared to the signal detected for the other member of the pair. As a second example, three of the four different nucleotide types may be detected under certain conditions, while the fourth nucleotide type may have no detectable label or be minimally detected under those conditions (e.g., minimal detection due to background fluorescence, etc.). Incorporation of the first three nucleotide types into a nucleic acid may be determined based on the presence of their corresponding signals, and incorporation of the fourth nucleotide type into a nucleic acid may be determined based on the absence or minimal detection of any signal. As a third example, one nucleotide type can include a label that is detected in two different channels, while the other nucleotide type is detected in no more than one of the channels. The three exemplary configurations above are not considered mutually exclusive and can be used in various combinations.An exemplary embodiment combining all three examples is a fluorescence-based SBS method that uses a first nucleotide type that is detected in a first channel (e.g., dATP having a label that is detected in the first channel when excited by a first excitation wavelength), a second nucleotide type that is detected in a second channel (e.g., dCTP having a label that is detected in the second channel when excited by a second excitation wavelength), a third nucleotide type that is detected in both the first and second channels (e.g., dTTP having at least one label that is detected in both channels when excited by the first and / or second excitation wavelength), and an unlabeled fourth nucleotide type that is not detected or is minimally detected in either channel (e.g., unlabeled dGTP).

[0210] Furthermore, as described in incorporated U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single channel. In such so-called single-dye sequencing methods, a first nucleotide type is labeled but the label is removed after the first image is generated, and a second nucleotide type is labeled only after the first image is generated. A third nucleotide type retains its label in both the first and second images, and a fourth nucleotide type remains unlabeled in both images.

[0211] Some embodiments may use sequencing by ligation techniques. Such techniques use DNA ligase to incorporate oligonucleotides and identify their incorporation. The oligonucleotides typically have different labels that correlate with the identity of specific nucleotides in the sequence to which the oligonucleotides hybridize. As with other SBS methods, images can be obtained after treating an array of nucleic acid sequences with labeled sequencing reagents. Each image shows nucleic acid features that incorporate a specific type of label. Because the sequence content of each feature varies, different features may or may not be present in different images, but the relative positions of the features remain constant within the image. Images obtained from ligation-based sequencing methods can be stored, processed, and analyzed as described herein. Exemplary SBS systems and methods that can be utilized with the methods and systems described herein are described in U.S. Patent Nos. 6,969,488, 6,172,218, and 6,306,597.

[0212] Some embodiments can use nanopore sequencing (Deamer, DW and Akeson, M., "Nanopores and Nucleic Acids: Perspectives on Ultrafast Sequencing," Trends Biotechnol., 18, 147-151 (2000); Deamer, D. and D. Branton, "Characterization of Nucleic Acids by Nanopore Analysis," Acc. Chem. Res., 35:817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin, and J. A. Golovchenko, "DNA Molecules and Organization in Solid-State Nanopore Microscopy," Nat. Mater, 2:611-615 (2003)). In such embodiments, the triple-indexed fragment passes through the nanopore. The nanopore can be a synthetic pore or a biological membrane protein, such as α-hemolysin. As the triple-index fragment passes through the nanopore, each base pair can be distinguished by measuring the fluctuation in the electrical conductance of the pore. (U.S. Pat. No. 7,001,792; Soni, G.V. and Meller, "Progress Towards Ultrafast DNA Sequencing Using Solid-Phase Nanopores," Clin. Chem. 53, 1996-2001 (2007); Healy, K., "Nanopore-Based Single-Molecule DNA Analysis," Nanomed., 2, 459-481 (2007); Cockroft, S.L., Chu, J., Amorin, M. and Ghadiri, M.R., "Single-Molecule Nanopore Device Detects DNA Polymerase Activity with Single-Nucleotide Resolution," J. Am Chem. Soc. 130, 818-820 (2008). Data obtained from nanopore sequencing can be stored, processed, and analyzed as described herein. In particular, the data can be processed as images according to the exemplary processing of optical and other images described herein.

[0213] Some embodiments can use methods involving real-time monitoring of DNA polymerase activity. Nucleotide incorporation can be detected via fluorescence resonance energy transfer (FRET) interactions between a fluorophore-containing polymerase and a γ-phosphate-labeled nucleotide, as described, for example, in U.S. Patent Nos. 7,329,492 and 7,211,414, or nucleotide incorporation can be detected using zero-mode waveguides, as described, for example, in U.S. Patent No. 7,315,019, and fluorescent nucleotide analogs and engineered polymerases, as described, for example, in U.S. Patent No. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082. Illumination can be restricted to a zeptoliter-scale volume around the surface-tethered polymerase so that incorporation of fluorescently labeled nucleotides can be observed with low background (Levene, MJ et al., "Zero-mode Waveguides for Single-Molecule Analysis at High Concentration," Science, 299, 682-686 (2003); Lundquist, PM et al., "Parallel Confocal Detection of Single Molecules in Real Time," Opt. Lett., 33, 1026-1028 (2008); Korlach, J. et al., "Selective Aluminum Passivation for Targeted Immobilization of Single DNA Polymerase Molecules in Zero-mode Waveguide Nanostructures," Proc. Natl. Acad. Sci. USA, 105, 1176-1181 (2008)). Images obtained from such methods can be stored, processed, and analyzed as described herein.

[0214] Some SBS embodiments involve the detection of protons released upon incorporation of a nucleotide into an extension product. For example, sequencing based on the detection of released protons can use commercially available electrical detectors and related technology from Ion Torrent (Guilford, Connecticut, a Life Technologies subsidiary), or the sequencing methods and systems described in U.S. Patent Application Publication Nos. 2009 / 0026082, 2009 / 0127589, 2010 / 0137143, and 2010 / 0282617. The methods described herein for amplifying target nucleic acids using equilibrium exclusion can be easily adapted to substrates used to detect protons. More specifically, the methods described herein can be used to generate clonal populations of amplicons used to detect protons.

[0215] The SBS method described above can be advantageously performed in a multiplex format, allowing multiple different triple index fragments to be manipulated simultaneously. In certain embodiments, the different triple index fragments can be processed in a common reaction vessel or on a specific substrate surface. This allows for convenient delivery of sequencing reagents, removal of unreacted reagents, and multiplexed detection of incorporation events. In embodiments using surface-bound target nucleic acids, the triple index fragments can be in an array format. In an array format, the triple index fragments can typically be bound to the surface in a spatially distinguishable manner. The triple index fragments can be bound by direct covalent binding, attachment to beads or other particles, or binding to a surface-bound polymerase or other molecule. The array can contain a single copy of the triple index fragment at each site (also called a feature), or multiple copies with the same sequence can be present at each site or feature. Multiple copies can be produced by amplification methods such as bridge amplification or emulsion PCR, as described in further detail herein.

[0216] The methods described herein can be used to fabricate, for example, at least about 10 features / cm 2, 100 features / cm 2 , 500 features / cm 2 , 1,000 features / cm 2 , 5,000 features / cm 2 , 10,000 features / cm 2 , 50,000 features / cm 2 , 100,000 features / cm 2 , 1,000,000 features / cm 2 , 5,000,000 features / cm 2 Arrays having features of any of a variety of densities, including 1000 sq. ft., ... or more, can be used.

[0217] The advantage of the method described herein is that it allows for multiple cm 2The present disclosure provides an integrated system capable of preparing and detecting nucleic acids using techniques known in the art, such as those exemplified herein. Accordingly, the integrated system of the present disclosure can include fluidic components capable of delivering amplification and / or sequencing reagents to one or more immobilized triplex index fragments, and the system includes components such as pumps, valves, reservoirs, and fluid lines. A flow cell can be configured and / or used in the integrated system for detecting target nucleic acids. Exemplary flow cells are described, for example, in U.S. Pat. Nos. 8,241,573 and 8,951,781. As exemplified for the flow cell, one or more of the fluidic components of the integrated system can be used in the amplification and detection methods. Taking the nucleic acid sequencing embodiment as an example, one or more of the fluidic components of the integrated system can be used to deliver sequencing reagents in the amplification methods described herein and the sequencing methods exemplified above. Alternatively, the integrated system can include separate fluidic systems for performing the amplification method and the detection method. Examples of integrated sequencing systems that can generate amplified nucleic acids and determine the sequence of nucleic acids include, but are not limited to, the MiSeq™ platform (Illumina, San Diego, California) and the device described in U.S. Pat. No. 8,951,781.

[0218] composition

[0219] Compositions are also provided herein. Various compositions can be obtained during the practice of the methods described herein. For example, compositions can be obtained that contain cells or nuclei with non-specifically or specifically bound hashing oligos. Other compositions include those that contain multiple populations of normalizing oligos as described herein. Compositions can be obtained that contain multiple compartments, each compartment containing multiple populations of normalizing oligos, or that contain multiple nuclei or cells, where the nuclei or cells are associated with multiple populations of normalizing oligos.

[0220] Illustrative Embodiments

[0221] Embodiment 1. A method for preparing a sequencing library comprising nucleic acids from a plurality of single nuclei or single cells, comprising: (a) providing a plurality of cells in a first plurality of compartments; (b) exposing the plurality of cells in each compartment to predetermined conditions; (c) contacting isolated nuclei from cells of each compartment or cells of each compartment with hashing oligos, at least one copy of the hashing oligo is associated with the isolated nucleus or cell; the hashing oligo comprises a hashing index, a hashing index in each partition including an index sequence that is different from the index sequences in other partitions to generate hash kernels or hash cells; (d) combining hash kernels or hash cells from different partitions to generate pooled hash kernels or pooled hash cells.

[0222] Embodiment 2. The method of embodiment 1, further comprising exposing the cells or isolated nuclei to a cross-linking compound to fix the cells or nuclei.

[0223] Embodiment 3. The method of any one of embodiments 1 to 2, wherein the cross-linking compound comprises paraformaldehyde, formalin, or methanol.

[0224] Embodiment 4. The method of any one of embodiments 1 to 3, wherein the predetermined condition comprises exposure to a drug.

[0225] Embodiment 5. The method of any one of embodiments 1 to 4, wherein the agent comprises a protein, a non-ribosomal protein, a polyketide, an organic molecule, an inorganic molecule, an RNA or RNAi molecule, a carbohydrate, a glycoprotein, a nucleic acid, a drug, or a combination thereof.

[0226] Embodiment 6 The method of any one of embodiments 1 to 5, wherein the hashing oligo comprises a single-stranded nucleic acid.

[0227] Embodiment 7. The method of any one of embodiments 1 to 6, wherein the hashing oligo consists of a single-stranded nucleic acid.

[0228] Embodiment 8. The method of any one of embodiments 1 to 7, wherein the nucleic acid of the hashing oligo comprises DNA, RNA, or a combination thereof.

[0229] Embodiment 9. The method of any one of embodiments 1 to 8, wherein the hashing oligo comprises a domain that mediates specific binding of the hashing oligo to the surface of a cell or nucleus.

[0230] Embodiment 10. The method of any one of embodiments 1 to 9, wherein the domain comprises a ligand, an antibody, or an aptamer.

[0231] Embodiment 11. The method of any one of embodiments 1 to 10, wherein the association between the hashing oligo and the cell or isolated nucleus is non-specific.

[0232] Embodiment 12 The method of any one of embodiments 1 to 11, wherein the non-specific association between the hashing oligo and the cells or isolated nuclei is due to absorption.

[0233] Embodiment 13. The method of any one of embodiments 1 to 12, further comprising processing the pooled hashed cells or pooled hashed nuclei using a single-cell combinatorial indexing method to obtain a sequencing library comprising nucleic acids from a plurality of single nuclei, wherein the nucleic acids comprise a plurality of indexes.

[0234] Embodiment 14. The method of any one of embodiments 1 to 13, wherein the single-cell combinatorial indexing method is single-nucleus transcriptome sequencing, single-cell transcriptome sequencing, single-cell transcriptome and transposon-accessible chromatin sequencing, single-nucleus whole-genome sequencing, single-nucleus sequencing of transposon-accessible chromatin, sci-HiC, drug-seq, sci-CAR, sci-MET, sci-Crop, sci-perturb, or sci-Crispr.

[0235] Embodiment 15. The method comprises: (e) distributing a subset of the pooled hashed cells or hashed nuclei into a second plurality of compartments and contacting each subset with a reverse transcriptase or DNA polymerase and a primer, wherein the primer in each compartment comprises a first index sequence that differs from the first index sequences in the other compartments in that it generates indexed nuclei that comprise indexed nucleic acid fragments; (f) combining the indexed cells or indexed nuclei to generate pooled indexed cells or pooled indexed nuclei; (g) distributing a subset of the pooled indexed cells or pooled indexed nuclei into a third plurality of compartments and introducing a second index sequence into the indexed nucleic acid fragments to generate dual-indexed cells or dual-indexed nuclei comprising dual-indexed nucleic acid fragments, wherein the introducing step comprises ligation, primer extension, amplification, or transposition; (h) combining the dual-indexed cells or dual-indexed nuclei to generate pooled dual-indexed nuclei or cells; (i) distributing a subset of the doubly-indexed cells or pooled doubly-indexed nuclei into a fourth plurality of compartments and introducing a third index sequence into the doubly-indexed nucleic acid fragments to generate triply-indexed cells or triply-indexed nuclei comprising triply-indexed nucleic acid fragments, wherein the introducing step comprises ligation, primer extension, amplification, or transposition; 15. The method of any one of embodiments 1 to 14, further comprising: (j) generating a sequencing library comprising transcriptome nucleic acids from a plurality of single nuclei by combining the triply indexed fragments.

[0236] Embodiment 16. (g) The method of any one of embodiments 1 to 15, comprising contacting each subset with a transposome complex, wherein the transposome complex in each compartment comprises a transposome and a second index sequence under conditions suitable for ligation of the second index sequence to an end of an indexed nucleic acid fragment comprising the first index sequence to generate a doubly-indexed nucleus comprising the doubly-indexed nucleic acid fragment, wherein the second index sequence is different from the second index sequences in the other compartments.

[0237] Embodiment 17. The method of any one of embodiments 1 to 16, comprising: (i) contacting each subset with a primer comprising a third index sequence and a universal primer sequence, wherein the contacting comprises conditions suitable for amplification and incorporation of the third index sequence into the ends of the doubly-indexed nucleic acid fragments, and wherein the third index sequence is different from the third index sequences in the other compartments.

[0238] Embodiment 18. The method of any one of embodiments 1 to 17, wherein the compartment comprises a well or a droplet.

[0239] Embodiment 19. providing a surface comprising a plurality of amplification sites; the amplification site comprises at least two populations of linked single-stranded capture oligonucleotides having free 3' ends; 19. The method of any one of embodiments 1 to 18, further comprising contacting the surface comprising the amplification sites with the triply indexed fragments under conditions suitable to generate a plurality of amplification sites each comprising a clonal population of amplicons from individual fragments comprising the plurality of indexes.

[0240] Embodiment 20. A composition comprising the hash cells or hash kernels of embodiment 1.

[0241] Embodiment 21. A composition comprising the pooled hash cells or pooled hash nuclei of embodiment 1.

[0242] Embodiment 22. A multiwell plate, wherein a compartment of the multiwell plate comprises a composition according to any one of embodiments 20 or 21.

[0243] Embodiment 23. A multiwell plate according to embodiment 22, wherein the compartments of the multiwell plate contain between 50 and 100 million cells or nuclei.

[0244] Embodiment 24. A droplet comprising a composition according to any one of embodiments 20 or 21.

[0245] Embodiment 25. The droplet of embodiment 24, wherein the droplet contains 50 to 100 million cells or nuclei.

[0246] Embodiment 26. A method for preparing a sequencing library comprising nucleic acids from a plurality of single nuclei or single cells, comprising: (a) providing a first plurality of compartments containing isolated nuclei or cells and contacting the isolated nuclei or cells in each compartment with a hashing oligo; at least one copy of the hashing oligo is associated with the isolated nucleus or cell; the hashing oligo comprises a nucleic acid and a hashing index, a hashing index in each partition including an index sequence that is different from the index sequences in other partitions to generate hash kernels or hash cells; (b) combining hash kernels or hash cells of different partitions to generate pooled hash kernels or pooled hash cells.

[0247] Embodiment 27. A method for preparing a sequencing library comprising nucleic acids from a plurality of single nuclei or single cells, comprising: (a) providing a first plurality of compartments containing isolated nuclei or cells and contacting the isolated nuclei or cells in each compartment with a hashing oligo; at least one copy of the hashing oligo is associated with the isolated nucleus or cell by absorption; the hashing oligo comprises a nucleic acid and a hashing index, a hashing index in each partition including an index sequence that is different from the index sequences in other partitions to generate hash kernels or hash cells; (b) combining hash kernels or hash cells of different partitions to generate pooled hash kernels or pooled hash cells.

[0248] Embodiment 28. A method for preparing a sequencing library comprising nucleic acids from a plurality of nuclei or cells, comprising: (a) providing a plurality of compartments comprising nuclei or cells, the nuclei or cells comprising hashing oligos comprising compartment-specific indices; (b) combining nuclei or cells from a different compartment into a second compartment to produce pooled hashed nuclei or pooled hashed cells.

[0249] Embodiment 29. The method of any one of embodiments 26 to 28, further comprising exposing cells of each compartment to a predetermined condition, or exposing cells of each compartment to a predetermined condition and then isolating nuclei from the plurality of cells prior to step (a).

[0250] Embodiment 30. The method of any one of embodiments 28 to 29, wherein the predetermined condition comprises exposure to a drug.

[0251] Embodiment 31. The method of any one of embodiments 28 to 30, wherein the agent comprises a protein, a non-ribosomal protein, a polyketide, an organic molecule, an inorganic molecule, an RNA or RNAi molecule, a carbohydrate, a glycoprotein, a nucleic acid, a drug, or a combination thereof.

[0252] Embodiment 32. A composition comprising multiple populations of normalization oligos, the composition comprising: a first population of normalization oligos comprising a first index sequence; and other populations of normalization oligos, each comprising a unique index sequence that is different from the index sequences of the other populations, wherein the concentration of each population is the same.

[0253] Embodiment 33. A composition comprising multiple populations of normalization oligos, the composition comprising a first population of normalization oligos comprising a first index sequence and other populations of normalization oligos, each comprising a unique index sequence that is different from the index sequences of the other populations, wherein the concentrations of at least two of the populations are different.

[0254] Embodiment 34. A composition comprising multiple populations of normalization oligos, the composition comprising a first population of normalization oligos comprising a first set of index sequences and other populations of normalization oligos, each comprising a unique set of index sequences that is different from the sets of index sequences of the other populations, wherein the concentration of each population is the same.

[0255] Embodiment 35. A composition comprising multiple populations of normalization oligos, the composition comprising a first population of normalization oligos comprising a first set of index sequences and other populations of normalization oligos, each comprising a unique set of index sequences that is different from the sets of index sequences of the other populations, wherein the concentrations of at least two of the populations are different.

[0256] Embodiment 36. The composition of any one of embodiments 32 to 35, wherein the composition comprises a population of 2 to 100 normalization oligos.

[0257] Embodiment 37. The composition of any one of embodiments 32 to 36, wherein the normalization oligo comprises single-stranded DNA.

[0258] Embodiment 38. The composition of any one of embodiments 32 to 37, wherein the normalization oligo comprises a unique molecular identifier.

[0259] Embodiment 39. The composition of any one of embodiments 32 to 38, wherein the normalization oligo comprises a universal sequence.

[0260] Embodiment 40. The composition of any one of embodiments 32 to 39, wherein the normalization oligo comprises a non-nucleic acid component.

[0261] Embodiment 41 The composition of embodiments 32 to 40, wherein the non-nucleic acid component comprises a protein.

[0262] Embodiment 42 The composition of any one of embodiments 32 to 41, wherein the non-nucleic acid component comprises a protein.

[0263] Embodiment 43. The composition of any one of embodiments 32 to 42, wherein a first population is present in the composition at the lowest concentration of normalization oligo and one of the other populations is present in the composition at the highest concentration of normalization oligo, and the lowest and highest concentrations differ by a factor of 1 to 10,000.

[0264] Embodiment 44. A plurality of compartments, each compartment containing a composition according to any one of embodiments 32 to 43.

[0265] Embodiment 45. A plurality of compartments according to embodiment 44, wherein the compartments comprise wells or droplets.

[0266] Embodiment 46. A plurality of compartments according to any one of embodiments 44 to 45, wherein each compartment further comprises a nucleus or a cell, and wherein the plurality of populations of normalization oligos are associated with the nucleus or cell.

[0267] Embodiment 47. The plurality of compartments according to any one of embodiments 44 to 46, wherein the concentration of normalization oligo in each population is selected from at least 0.001 zeptomolar to no more than 100 attomoles.

[0268] Embodiment 48. A population of nuclei or cells, wherein the nuclei or cells comprise the composition of any one of embodiments 32 to 47, and wherein each population member of the normalization oligo is associated with a nucleus or cell.

[0269] Embodiment 49. The population of embodiment 48, wherein the association between the nuclei or cells and the normalization oligo is non-specific.

[0270] Embodiment 50. A method for normalizing a sequencing library comprising nucleic acids from a plurality of single nuclei or single cells, comprising: (a) providing a first plurality of compartments comprising isolated nuclei or cells; (b) contacting the isolated nuclei or cells of each compartment with the composition of any one of embodiments 32 to 35, wherein members of each population of normalization oligos are associated with the isolated nuclei or cells; (c) combining the labeled nuclei or labeled cells of the different compartments to generate pooled labeled nuclei or pooled labeled cells.

[0271] Embodiment 51. The method of embodiment 50, further comprising exposing the cells of each compartment to a predetermined condition, or exposing the cells of each compartment to a predetermined condition and then isolating nuclei from the plurality of cells prior to step (a).

[0272] Embodiment 52. The method of any one of embodiments 50 to 51, wherein the predetermined condition comprises exposure to a drug.

[0273] Embodiment 53. The method of any one of embodiments 50 to 52, wherein the agent comprises a protein, a non-ribosomal protein, a polyketide, an organic molecule, an inorganic molecule, an RNA or RNAi molecule, a carbohydrate, a glycoprotein, a nucleic acid, a drug, or a combination thereof.

[0274] Embodiment 54. Prior to step (b), contacting the isolated nuclei or cells of each compartment with a hashing oligo, at least one copy of the hashing oligo is associated with the isolated nucleus or cell; the hashing oligo comprises a nucleic acid and a hashing index, a hashing index within each compartment containing an index sequence that is different from the index sequences within other compartments and different from the index sequences of normalization oligos present within the compartment to generate labeled hash kernels or labeled hash cells; 54. The method of any one of embodiments 50 to 53, further comprising combining labeled hash kernels or labeled hash cells of different compartments to generate pooled labeled hash kernels or pooled labeled hash cells.

[0275] Embodiment 55. The method of any one of embodiments 50 to 54, further comprising exposing the normalized oligonucleotide cells or isolated nuclei to a cross-linking compound to fix the cells or nuclei.

[0276] Embodiment 56 The method of any one of embodiments 50 to 55, wherein the cross-linking compound comprises paraformaldehyde, formalin, or methanol.

[0277] Embodiment 57. The method of any one of embodiments 50 to 56, wherein the association between the normalization oligo and the cell or isolated nucleus is non-specific.

[0278] Embodiment 58. The method of any one of embodiments 50 to 57, wherein the non-specific association between the normalization oligo and the cells or isolated nuclei is due to absorption.

[0279] Embodiment 59. The method of any one of embodiments 50 to 58, further comprising processing the pooled labeled hashed cells or pooled labeled hashed nuclei using a single-cell combinatorial indexing method to obtain a sequencing library comprising nucleic acids from a plurality of single nuclei or single cells, wherein the nucleic acids comprise a plurality of indexes.

[0280] Embodiment 60. The method of any one of embodiments 50 to 59, wherein the single-cell combinatorial indexing method is single-nucleus transcriptome sequencing, single-cell transcriptome sequencing, single-cell transcriptome and transposon-accessible chromatin sequencing, single-nucleus whole-genome sequencing, single-nucleus sequencing of transposon-accessible chromatin, sci-HiC, drug-seq, sci-CAR, sci-MET, sci-Crop, sci-perturb, or sci-Crispr.

[0281] Embodiment 61. A method for normalizing a sequencing library comprising nucleic acids from a plurality of single nuclei or single cells, comprising: (a) providing an isolated nucleus or cell; (b) contacting the isolated nucleus or cell with the composition of any one of embodiments 32 to 35, wherein a member of each population of normalization oligos is associated with the isolated nucleus or cell; (c) distributing a subset of labeled nuclei or labeled cells into a plurality of compartments.

[0282] Embodiment 62. The method of embodiment 61, further comprising exposing the cells to a predetermined condition, or exposing the cells to a predetermined condition and then isolating nuclei from the plurality of cells prior to step (a).

[0283] Embodiment 63. The method of any one of embodiments 61 to 62, wherein the predetermined condition comprises exposure to a drug.

[0284] Embodiment 64. The method of any one of embodiments 61 to 63, wherein the agent comprises a protein, a non-ribosomal protein, a polyketide, an organic molecule, an inorganic molecule, an RNA or RNAi molecule, a carbohydrate, a glycoprotein, a nucleic acid, a drug, or a combination thereof.

[0285] Embodiment 65. After step (c), contacting the isolated nuclei or cells of each compartment with a hashing oligo, at least one copy of the hashing oligo is associated with the isolated nucleus or cell; the hashing oligo comprises a nucleic acid and a hashing index, a hashing index within each compartment containing an index sequence that is different from the index sequences within other compartments and different from the index sequences of normalization oligos present within the compartment to generate labeled hash kernels or labeled hash cells; 65. The method of any one of embodiments 61 to 64, further comprising combining labeled hash kernels or labeled hash cells of different compartments to generate pooled labeled hash kernels or pooled labeled hash cells.

[0286] Embodiment 66 The method of any one of embodiments 61 to 65, further comprising exposing the normalized oligonucleotide cells or isolated nuclei to a cross-linking compound to fix the cells or nuclei.

[0287] Embodiment 67. The method of any one of embodiments 61 to 66, wherein the cross-linking compound comprises paraformaldehyde, formalin, or methanol.

[0288] Embodiment 68. The method of any one of embodiments 61 to 67, wherein the association between the normalization oligo and the cell or isolated nucleus is non-specific.

[0289] Embodiment 69. The method of any one of embodiments 61 to 68, wherein the non-specific association between the normalization oligo and the cells or isolated nuclei is due to absorption.

[0290] Embodiment 70. The method of any one of embodiments 61 to 69, further comprising processing the pooled labeled hashed cells or pooled labeled hashed nuclei using a single-cell combinatorial indexing method to obtain a sequencing library comprising nucleic acids from a plurality of single nuclei or single cells, wherein the nucleic acids comprise a plurality of indexes.

[0291] Embodiment 71. The method of any one of embodiments 61 to 70, wherein the single-cell combinatorial indexing method is single-nucleus transcriptome sequencing, single-cell transcriptome sequencing, single-cell transcriptome and transposon-accessible chromatin sequencing, single-nucleus whole genome sequencing, single-nucleus sequencing of transposon-accessible chromatin, sci-HiC, drug-seq, sci-CAR, sci-MET, sci-Crop, sci-perturb, or sci-Crispr.

[0292] [Example]

[0293] Example 1

[0294] Massively multiplexed chemical transcriptomics at single-cell resolution

[0295] High-throughput chemical screens typically use crude assays such as cell viability, limiting what can be learned about mechanisms of action, off-target effects, and heterogeneous responses. Here, we introduce "sci-Plex," which uses "nuclear hashing" to quantify global transcriptional responses to thousands of independent perturbations at single-cell resolution. As a proof-of-concept, we applied sci-Plex to screen three cancer cell lines exposed to 188 compounds. In total, we profiled approximately 650,000 single-cell transcriptomes across approximately 5,000 independent samples in a single experiment. Our results reveal substantial cell-to-cell heterogeneity in response to specific compounds, commonalities across families of compounds, and insights into distinct properties within families. Specifically, our results with histone deacetylase inhibitors support the view that chromatin functions as an important reservoir of acetate in cancer cells. This example is also available as Srivatsan et al., 2020, Science, 367:45-51.

[0296] To enable cost-effective high-throughput screening (HTS) with single-cell transcriptome sequencing (scRNAseq)-based phenotyping, we describe a novel sample labeling (hashing) strategy that relies on labeling nuclei with unmodified single-stranded DNA oligos. Recent improvements in single-cell combinatorial indexing (sciRNA-seq3) have reduced the cost of scRNAseq library preparation to less than $0.01 per cell, allowing millions of cells to be profiled per experiment (21). Here, we combine nuclear hashing and sciRNA-seq into a single multiplexed transcriptomics workflow in a process we call "sci-Plex." As proof-of-concept, we perform HTS in three cancer cell lines using sci-Plex, profiling thousands of independent perturbations in a single experiment. We further explore how chemical transcriptomics at single-cell resolution can shed light on mechanisms of action. Most notably, we found that changes in gene regulation as a result of treatment with histone deacetylase (HDAC) inhibitors are consistent with a model in which the inhibitors impede proliferation by limiting the ability of cells to withdraw acetate from chromatin ( 22 , 23 ).

[0297] result

[0298] Nuclear hashing enables multi-sample sci-RNA-seq

[0299] Single-cell combinatorial indexing (sci-) methods use split-pool barcodes to specifically label the molecular weights of large numbers of single cells or nuclei. (24) Samples can be barcoded with these same indexes, for example, by placing each sample in its own well during reverse transcription in sci-RNA-seq. (21, 25) However, such enzymatic labeling on the scale of thousands of samples is operationally infeasible and cost-prohibitive. To enable single-cell molecular profiling of large numbers of independent samples within a single sci-experiment, we set out to develop a low-cost labeling procedure.

[0300] We noticed that single-stranded DNA (ssDNA) specifically stained the nuclei of permeabilized cells but not intact cells (Figures 4A and 5A). Therefore, we hypothesized that polyadenylated ssDNA oligonucleotides could be used to label a population of nuclei in a manner compatible with sci-RNA-seq (Figures 4B and 5B). To test this concept, we performed a "barnyard" experiment. Human (HEK293T) and mouse (NIH3T3) cells were separately seeded into 48 wells of a 96-well culture plate. Next, nuclear lysis was performed in the presence of specific polyadenylated ssDNA oligos ("hash oligos") in each well, and the resulting nuclear suspension was fixed with paraformaldehyde. After labeling, or "hashing," the nuclei with molecular barcodes, the nuclei were pooled and a two-level sci-RNA-seq experiment was performed. Because the hash oligos are polyadenylated, they could be ligated and indexed in the same way as endogenous mRNAs. As intended, we recovered reads corresponding to both endogenous mRNAs [median 4740 unique molecular identifiers (UMIs) per cell] and hashed oligos (median 270 UMIs per cell).

[0301] We devised a statistical framework to identify hashed oligos associated with each cell at frequencies above background (Table S1). We observed 99.1% discrepancy between species assignments based on hashed oligos versus the endogenous cellular transcriptome (Figure 4C and Figures 5C-5F). Furthermore, the association of hashed oligos with nuclei was stable to freeze-thaw cycles, highlighting the opportunity to label and preserve samples (Figure 4D and Figures 5G-5H). These results demonstrate that hashed oligos stably label nuclei in a manner compatible with sci-RNA-seq.

[0302] [Table 1]

[0303] In sci- experiments, a "collision" occurs when two or more cells are labeled by chance with the same combination of barcodes (24). To evaluate hashing as a means of detecting doublets resulting from collisions, we varied the number of nuclei loaded per polymerase chain reaction well, resulting in a range of predicted collision rates (7 to 23%) that were in good agreement with observations (Figure 5I). Hash oligos facilitated identification of the majority of interspecies doublets (95.5%) and intraspecies doublets that would otherwise be undetectable (Figure 4E and Figure 5J and Figure 5K).

[0304] sci-Plex enables chemical transcriptomics multiplexing at single-cell resolution.

[0305] Next, we evaluated whether nuclear hashing could enable chemical screening by labeling cells that underwent specific perturbations, followed by single-cell transcriptional profiling as a high-content phenotypic assay. We exposed A549, a human lung adenocarcinoma cell line, to one of four compounds: dexamethasone (a corticosteroid agonist), Nutlin-3a (a p53-Mdm2 antagonist), BMS-345541 (an inhibitor of nuclear factor kB-dependent transcription), or vorinostat (suberoylanilide hydroxamic acid (SAHA), an HDAC inhibitor) in triplicate at seven doses over 24 hours, for a total of 84 drug-dose-duplicate combinations, plus an additional vehicle control (Figures 6A and 7A). We labeled nuclei from each well and subjected them to sci-RNA-seq2 (Figures 7B-7D and Table 1).

[0306] We used Monocle 3 (21) to visualize these data using Uniform Manifold Approximation and Projection (26) (UMAP) and Louvain community detection to identify compound-specific clusters of cells distributed in a dose-dependent manner (Figures 6B and 6C, and 7E and 7F). To quantify the "population-average" transcriptional response of A549 cells to each of the four drugs, we modeled the expression of each gene as a function of dose via general linear regression. A total of 7,561 genes were sensitive to at least one drug, and 3,189 genes were differentially expressed in response to multiple drugs (Figure 8A, data not shown). These included canonical targets of dexamethasone (Figure 6D) and Nutlin-3a (Figure 6E). Gene ontology analysis of differentially expressed genes revealed the involvement of drug-specific pathways (e.g., hormone signaling for dexamethasone, p53 signaling for Nutlin-3a, Figure 8B). Additionally, we assessed whether the number of cells recovered at each concentration could be used to infer toxicity similar to that observed in conventional screens. After fitting a response curve to the number of recovered cells, we inferred a "viability score" from the sci-Plex data, a metric consistent with the "gold standard" measurement (Figures 6F and 7G-I).

[0307] sci-Plex scales to thousands of samples, enabling HTS.

[0308] To assess the scaling of sci-Plex for HTS, we screened 188 compounds targeting a diverse range of enzymes and molecular pathways (Figure 9A). Half of this panel was selected to target transcriptional and epigenetic regulators. The other half was chosen to sample different mechanisms of action. We exposed three well-characterized human cancer cell lines, A549 (lung adenocarcinoma), K562 (chronic myeloid leukemia), and MCF7 (breast adenocarcinoma), to each of these 188 compounds at four doses (10 nM, 100 nM, 1 mM, and 10 mM) twice, randomizing compound and dose across well positions in replicate culture plates (data not shown). These conditions, along with vehicle controls, accounted for 4608 of the 4992 cell populations independently treated in this experiment. After treatment, we lysed the cells to expose the nuclei, hashed them with specific combinations of two oligos (Figure 10A), and performed sciRNA-seq3 (21). After sequencing and filtering based on hash purity (Figures 10B–10F), we obtained transcriptomes of 649,340 single cells with median mRNA UMI counts of 1271, 1071, and 2407 for A549, K562, and MCF7, respectively (Figure 11A). The aggregate expression profiles of each cell type were highly concordant between replicate wells (Pearson correlation = 0.99) (Figure 11B).

[0309] Separate visualization of sci-RNA-seq profiles for each cell line revealed compound-specific transcriptional responses and patterns common to multiple compounds. For each cell line, UMAP projected most cells within a central cluster adjacent to the smallest cluster (Figure 9B). These smaller clusters were primarily composed of cells treated with compounds from one or two compound classes (Figures 12 and 13A-C). For example, A549 cells treated with the synthetic glucocorticoid receptor agonist triamcinolone acetonide were significantly enriched in one such small cluster, comprising 95% of the cells [Fisher's exact test, false discovery rate (FDR) <1%, Figures 13D-E]. While many drugs were associated with seemingly homogeneous transcriptional responses, we also identified cases in which distinct transcriptional states were induced by the same drug. For example, in A549, the microtubule-stabilizing compounds epothilone A and epothilone B were associated with three such foci composed of cells from both compounds at four doses each (Figures 13F and 13G). Cells within each focus were distinct from each other but transcriptionally similar to other treatments, namely, the recently identified microtubule-destabilizing agent rigosertib (27), the SETD8 inhibitor UNC0397, or untreated proliferating cells (Figure 13H).

[0310] Next, we assessed the effect of each drug on the "population-average" transcriptome of each cell line. In total, 6,238 genes were differentially expressed in a dose-dependent manner in at least one cell line (FDR < 5%, Figure 14, data not shown). Bulk RNA-seq measurements collected for five compounds at four doses and vehicle were consistent with average gene expression values ​​and estimated effect sizes across identically treated single cells, although correlations between small effect sizes were reduced (Figure 15). Furthermore, sci-Plex dose-dependent effect profiles correlated with compound-matched L1000 measurements (11) (Figure 16).

[0311] Cell cycle-related genes varied widely between individual cells, and many drugs reduced the fraction of cells expressing proliferation marker genes (Figures 17 and 18). In principle, scRNA-seq should be able to distinguish changes in the fraction of cells in different transcriptional states from changes in gene regulation within those states. In contrast, bulk transcriptome profiling confounds these two signals (Figure 19A) (14). Therefore, we tested for dose-dependent differential expression in subsets of cells expressing high versus low levels of proliferation marker genes in response to the same drug (Figure 19B). The correlation between dose-dependent effects on the two fractions of each cell type varied between drug classes (Figure 19C), with some overtly discordant effects for individual compounds (Figure 19D). Viability analysis, performed similarly to the pilot experiments, revealed that only 52 (27%) compounds caused a 50% or greater reduction in viability after drug exposure at the highest dose (Figures 9C and 11C). Among the drugs that reduced viability, we observed enhanced sensitivity of K562 to the Src and Abl inhibitor bosutinib (Fig. 9C), a result confirmed by cell counts (Fig. 20A). This result is consistent with the observed increased sensitivity of K562 cells, which contain a constitutively active BCR-ABL fusion kinase (28), and hematopoietic and lymphoid cancer cell lines to Abl inhibitors (29) (Fig. 20B).

[0312] To assess whether each compound induced similar responses in the three cell lines, we clustered the compounds using the dose-dependent gene effect sizes as loadings for each cell line (Figures 21-24). Joint analysis of the three cell lines revealed common and cell-type-specific responses to various compounds (Figures 25 and 26). For example, trametinib, a mitogen-activated protein kinase kinase (MEK) inhibitor, induced transcriptionally distinct responses in MCF7 cells. Examination of UMAP projections revealed that trametinib-treated MCF7 cells were scattered among vehicle controls, reflecting a limited effect. In contrast, trametinib-treated A549 and K562 cells, which contain activating KRAS and ABL mutations, respectively (30), clustered tightly, consistent with a robust and specific transcriptional response to inhibition of MEK signaling by trametinib (Figure 9D). Furthermore, we observed that these A549 and K562 cells appeared proximal to clusters enriched with inhibitors of HSP90, a key chaperone for protein folding (Figure 9D). This observation was supported by concordant changes in HSP90AA1 expression in trametinib-treated cells (Figure 9E). Analysis of connectivity map data (11, 12) revealed further evidence that MEK inhibitors indeed induce gene expression signatures highly similar to HSP90 perturbation, particularly in A549 (Figure 20C), but not in MCF7 (Figures 20D and 20E). These results are consistent with previous observations of HSP90AA1 regulation downstream of MEK signaling (31) and suggest that similarities in single-cell transcriptomes treated with distinct compounds can highlight drugs targeting convergent molecular pathways.

[0313] Deducing the chemical and mechanistic properties of HDAC inhibitors

[0314] For each of the three cell lines, the most significant compound responses consisted of cells treated with one of 17 HDAC inhibitors (Figure 9B, dark blue; data not shown). To assess the similarity of dose-response trajectories between cell lines, we used a mutual nearest neighbor (MNN) matching approach (32) to align HDAC-treated and vehicle-treated cells from all three cell lines and generate consensus HDAC inhibitor trajectories, termed "pseudo-dose" [similar to "pseudo-time" (33)] (Figures 27A and 28). We observed that some HDAC inhibitors induced homogeneous responses, with nearly all cells localized to a relatively narrow range of HDAC inhibitor trajectories at each dose (e.g., pranostin in A549), whereas other drugs induced much greater cellular heterogeneity (Figures 27B and 29).

[0315] Such heterogeneity can be explained by cells asynchronously executing a defined transcriptional program, and the dose of drug to which cells are exposed modulates the rate at which cells progress through the program. To test this hypothesis, we sequenced the transcriptomes of 64,440 A549 cells treated for 72 hours with one of 48 compounds, including many of the HDAC inhibitors from a large-scale sci-Plex screen. Taking into account confluency-dependent cell cycle effects and MNN alignment (Figures 30 and 31), co-embedded UMAP projections revealed new focal concentrations of cells at 72 hours that were not apparent at the 24 hour time point, e.g., SRT1024 (Figure 32). However, for the majority of HDAC inhibitors tested, we did not observe cells at a given dose moving further along aligned HDAC trajectories at 72 hours (Figure 33). This suggests that the dose of many HDAC inhibitors governs the magnitude of the cellular response rather than the rate at which cells progress, and that the observed heterogeneity is not solely due to asynchrony (Figure 33).

[0316] Next, we assessed whether the target affinity of a given HDAC inhibitor explains its overall transcriptional response to the compound. A dose-response model was used to calculate the transcriptional median effective concentration (TC) for each compound. 50), i.e., the concentration required to drive cells midway through a simulated HDAC inhibitor dosing trajectory, was estimated (Figure 34A, data not shown). To compare transcriptionally derived potency measures with each compound's biochemical properties, we compared the published median inhibitory concentration (IC) of each compound from in vitro assays performed on eight purified HDAC isoforms. 50 ) were collected (data not shown). With the exception of two relatively insoluble compounds, our calculated TC50 Values ​​increased as a function of compound IC50 values ​​(Figures 27C and 34B and 34C).

[0317] To assess the components of the HDAC inhibitor trajectory, we performed differential expression analysis using sham treatment as a continuous covariate. Of the 4,308 genes that were significantly differentially expressed on this consensus trajectory, 2,081 (48%) responded in a cell-type-dependent manner, and 942 (22%) showed the same pattern across all three cell lines (Figures 35A and 35B, data not shown). One notable pattern shared by the three cell lines was the enrichment of genes and pathways indicative of progression toward cell cycle arrest (Figures 35C and 36A and 36B). DNA content staining and flow cytometry confirmed that HDAC inhibition resulted in the accumulation of cells in the G2 / M phase of the cell cycle (34) (Figures 36C and 36D).

[0318] Shared responses to HDAC inhibition included not only cell cycle arrest but also altered expression of genes involved in cellular metabolism (Figure 35C). Histone acetyltransferases and deacetylases regulate chromatin accessibility and transcription factor activity through the addition or removal of charged acetyl groups (35-37). Acetate, a product of histone deacetylation via HDAC classes I, II, and IV and a precursor of acetyl-CoA, is required for histone acetylation but also plays an important role in metabolic homeostasis (23, 38, 39). Inhibiting nuclear deacetylation limits the recycling of chromatin-bound acetyl groups for both catabolic and anabolic processes (39). Accordingly, we observed that HDAC inhibition sequestered acetate, resulting in a significant increase in acetylated lysine levels after exposure to 10 mM doses of the HDAC inhibitors pracinostat and abexinostat (Figure 37).

[0319] Further examination of the pseudo-dose-dependent genes revealed that enzymes important for cytosolic acetyl-CoA synthesis from either citrate (ACLY) or acetate (ACSS2) were upregulated (Figure 38A). Genes involved in cytosolic citrate homeostasis (GLS, IDH1, and ACO1), cellular citrate import (SLC13A3), and mitochondrial citrate production and export (CS, SLC25A1) were also upregulated. Upregulation of SIRT2, which deacetylates tubulin, was also observed in response to HDAC inhibition.

[0320] These transcriptional responses, along with the increase in chromatin-bound acetate, suggest a metabolically consequential depletion of cellular acetyl-CoA reserves in HDAC-inhibited cells (Figure 38B). To further validate this, we attempted to shift the distribution of cells along the HDAC inhibitor trajectory by modulating cellular acetyl-CoA levels. A549 and MCF7 cells were treated with pracinostat in the presence and absence of acetyl-CoA precursors (acetate, pyruvate, or citrate) or enzyme inhibitors involved in replenishing the acetyl-CoA pool (ACLY, ACSS2, or PDH). After treatment, cells were harvested and processed using sci-Plex to generate trajectories for each cell line (Figures 39 and 40). In both A549 and MCF7 cells, acetate, pyruvate, and citrate supplementation prevented pracinostat-treated cells from reaching the end of the HDAC inhibitor trajectory (Figures 39F, 39J, 39H, and 39L). In MCF7 cells, both ACLY and ACSS2 inhibition shifted cells further along the HDAC inhibitor trajectory, whereas no such shift was observed in A549 cells (Figures 39G, 39K, 39I, and 39M). ​​Together, these results suggest that a key feature of the cellular response to HDAC inhibitors, and in some cases their associated toxicity, is the induction of an acetyl-CoA-deficient state.

[0321] Discussion

[0322] We introduce sci-Plex, a large-scale multiplexed platform for single-cell transcriptomics. sci-Plex uses chemical fixation to cost-effectively and irreversibly label nuclei with short, unmodified ssDNA oligos. In the proof-of-concept experiments described herein, we applied sci-Plex to quantify the dose-dependent responses of cancer cells to 188 compounds through both high-content (global transcriptome) and high-resolution (single-cell) assays. By profiling several different cancer cell lines, we distinguished between shared and cell-line-specific molecular responses to each compound.

[0323] sci-Plex offers several distinct advantages over traditional HTS. It can distinguish distinct effects of compounds on cell subsets (including complexes in complex in vitro systems such as cell reprogramming, organoids, and synthetic embryos), mask heterogeneity in cellular responses to perturbations, and measure how drugs shift the relative proportions of transcriptionally distinct cell subsets. Highlighting these features, this study provides insight into the mechanism of action of HDAC inhibitors. Specifically, we found that the primary transcriptional response to HDAC inhibitors involves significant shifts in genes related to cell cycle arrest and acetyl-CoA metabolism. For some HDAC inhibitors, distinct heterogeneity was observed in the observed response at the single-cell level. While HDAC inhibition is traditionally thought to act through mechanisms directly involved in chromatin regulation, our data support an alternative, non-mutually exclusive, model in which HDAC inhibitors impair growth and proliferation by interfering with cancer cells' ability to withdraw acetate from chromatin (22, 23, 39). Therefore, variation in cellular acetate reservoirs is a potential explanation for their heterogeneous response to HDAC inhibitors.

[0324] As the cost of single-cell sequencing continues to fall, there may be considerable opportunities to leverage sci-Plex for fundamental and applied goals in biomedicine. The proof-of-concept experiment described here, consisting of nearly 5,000 independent treatments with transcriptional profiling of over 100 single cells per treatment, could potentially be expanded toward a comprehensive, high-resolution atlas of cellular responses to pharmacological perturbations (e.g., hundreds of cell lines or genetic backgrounds, thousands of compounds, multi-channel single-cell profiling, etc.). The ease and low cost of oligohashing, coupled with the flexibility and exponential scalability of single-cell combinatorial indexing, will facilitate this goal.

[0325] quotation

[0326] 1.JRBroach, J.Thorner, Nature 384(Lecture), 14-16(1996).

[0327] 2.DAPereira, JWilliams, Br.J.Pharmacol.152,53-61(2007).

[0328] 3.D.Shum, J. Enzyme Inhib.Med.Chem.23, 931–945(2008).

[0329] 4.C.Yuら, Nat.Biotechnol.34, 419–423(2016).

[0330] 5. ZEPerlman, Science 306, 1194-1198 (2004).

[0331] 6. Y. Futamura, Chem.Biol 19, 1620–1630 (2012).

[0332] 7.J.Kangら, Nat.Biotechnol.34, 70-77(2016).

[0333] 8.KLHuss,PEBlonigen,RMCampbell,J.Biomol.Screen.12,578-584(2007).

[0334] 9.C.Yeら, Nat.Commun.9, 4307(2018).

[0335] 10.ECBush, Nat.Commun.8, 105(2017).

[0336] Cell 171, 1437–1452.e17(2017).

[0337] 12.J.Lamb, Science 313, 1929-1935 (2006).

[0338] 13.MB Elowitz, AJLevine, EDSigia, PSSwain, Science 297, 1183-1186 (2002).

[0339] 14.C.Trapnell, Genome Res.25, 1491–1498(2015).

[0340] 15. SMShaffer, Nature 546, 431-435 (2017).

[0341] 16. SLSpencer, S. Gaudet, JGAlbeck, JMBurke, PKSorger, Nature 459, 428-432 (2009).

[0342] 17.M.Stoeckius, Genome Biol.19, 224(2018).

[0343] 18. J. Gehring, JHPark, S. Chen, M. Thomson, L. Pachter, Highly multiplexed single-cell RNA-seq for defining cell population and transcriptional spaces 、bioRxiv 315333[プレプリント]2018 May 5, doi.org / 10.1101 / 315333.

[0344] 19.CSMcGinnis, Nat.Methods 16, 619-626(2019).

[0345] 20. DWLee, JHLee, D.Bang, Sci.Adv.5, eaav2249(2019).

[0346] 21. J. Cao, Nature 566, 496-502(2019).

[0347] 22. MA McBrian et al., Mol. Cell 49, 310-321 (2013).

[0348] 23. SA Comerford et al., Cell 159, 1591-1602 (2014).

[0349] 24. DACusanovich et al., Science 348, 910-914 (2015).

[0350] 25. J. Cao et al., Science 357, 661-667 (2017).

[0351] 26. L. McInnes, J Healy, UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction.arXiv:1802.03426[stat.ML] (February 9, 2018).

[0352] 27. M. Jost et al., Mol. Cell 68, 210-223.e6 (2017).

[0353] 28. G. Grosveld et al., Mol. Cell. BiBiol 6, 607-616 (1986).

[0354] 29. EK Greuber, P. Smith-Pearson, J. Wang, AMPendergast, Nat. Rev. Cancer 13, 559-571 (2013).

[0355] 30. J. Barretina et al., Nature 483, 603-607 (2012).

[0356] 31. C. Dai et al., J. Clin. Invest. 122, 3742-3754 (2012).

[0357] 32. L. Haghverdi, ATLLun, MDMorgan, JC Marioni, Nat. Biotechnol. 36, 421-427 (2018).

[0358] 33. C. Trapnell, Nat. Biotechnol. 32, 381-386 (2014).

[0359] 34. W. Brazelle, PLOS ONE 5, e14335 (2010).

[0360] 35. J.-S Roe, F. Mercan, K. Rivera, DJ Pappin, CR Vakoc, Mol. Cell 58, 1028-1039 (2015).

[0361] 36.JE Brownell, Cell 84, 843-851 (1996).

[0362] 37. J. Taunton, CA Hassig, SL Schreiber, Science 272, 408-411 (1996).

[0363] 38. SKKurdiani, Curr. Opin. Genet. Dev. 26, 53-58 (2014).

[0364] 39. K.E. Wellen, Science 324, 1076-1080 (2009).

[0365] Materials and methods

[0366] Cell culture

[0367] A549 and K562 cells were kind gifts from Dr. Robert Bradley (University of Washington) and Dr. David Hawkins (University of Washington), respectively. MCF7 (catalog no. HTB-22), NIH3T3 (catalog no. CRL-1658), and HEK293T (catalog no. CRL-11268) cells were purchased from ATCC. A549 and MCF7 cells were cultured in DMEM (ThermoFisher, catalog no. 11995073) medium supplemented with 10% FBS (ThermoFisher, catalog no. 26140079) and 1% penicillin-streptomycin (ThermoFisher, catalog no. 15140122). K562 cells were cultured in RPMI 1640 (Fisher Scientific, Cat. No. 11-875-119) supplemented with 10% FBS and 1% penicillin-streptomycin and maintained at 0.2-1 x 10 cells / mL. All cells were cultured at 37°C with 5% CO. Adherent cells were split when they reached 90% confluence by washing with DPBS (Life Technologies, Cat. No. 14190-250) and maintained at 0.2-1 x 10 cells / mL using TryPLE (Fisher Scientific). The cells were trypsinized using a trypsinizer (Chemicals Scientific, Cat. No. 12-604-039) and split at 1:4 (MCF7) or 1:10 (A549, NIH3T3 and HEK293T).

[0368] Preparation of compounds

[0369] Dexamethasone was purchased from Sigma-Aldrich and resuspended in molecular biology-grade ethanol (Fisher Scientific). BMS-345541 (S8044), Vorinostat (S1047), and Nutlin-3a (S8059) were obtained from Selleck Chemicals and resuspended in DMSO (VWR Scientific, 97063-136). Cherry-pick 96-well compound screens were obtained from Selleck Chemicals and resuspended in 10 mM DMSO (data not shown). Compounds were diluted in their respective vehicles to 1000-fold their desired treatment concentrations and stored at -80°C until use.

[0370] Drug Treatment

[0371] For 96-well experiments, adherent cells were trypsinized, washed with PBS, and plated at 25,000 cells per well in 100 μL of medium in tissue culture-treated 96-well flat-bottom plates (Thermo Fisher Scientific, catalog number 12-656-66). Suspension cells were washed with PBS and plated at 25,000 cells per well in 100 μL of medium in 96-well V-bottom tissue culture plates (Thermo Fisher Scientific, catalog number 549935). Cells were allowed to recover for 24 h before being treated with 1 μL of a 1:10 dilution of the appropriate compound or vehicle in PBS, maintaining a 0.1% vehicle concentration in all wells. Cells were then exposed to the indicated concentrations of small molecules for 24 or 72 h. In experiments in which cells were co-treated with HDAC inhibitors and either acetate, pyruvate, citrate, ACSS2 inhibitor (EMD Millipore, catalog no. 533756), ACLY inhibitor (Cayman Chemicals, BMS-303141, catalog no. 943962-47-8), or PDH inhibitor (Cayman Chemicals, catalog no. 504817), cells were treated for 24 hours after plating and harvested 24 hours later. In this series of experiments, all wells contained a final concentration of 0.2% DMSO to accommodate treatment with both HDAC inhibitors and inhibitors of metabolic processes.

[0372] CellTiter Glo

[0373] A549, MCF7, and K562 cells were seeded in 96-well plates, allowed to attach for 24 hours, and then treated with BMS345541, dexamethasone, Nutlin-3A, or SAHA as described above. After 24 hours of treatment, plates were allowed to reach room temperature, and viability was estimated using the CellTiter-Glo Viability Assay (Promega) according to the manufacturer's instructions. Luminescence was recorded using a BioTek Synergy Plate Reader. Luminescence readings for each drug treatment were normalized to the mean luminescence intensity of vehicle DMSO-treated wells.

[0374] Cell counts of bosutinib-exposed cells

[0375] A549, MCF7, and K562 cells were seeded at 2.8 x 10 cells per well in 12-well plates. After 24 hours, A549 and MCF7 cells were allowed to adhere, and then exposed to 0.1, 1, or 10 μM bosutinib or DMSO vehicle control for 24 hours. After treatment, adherent cells were detached using TrypLE or directly resuspended in 1 mL of medium, and cells were counted using a Countess II FL automated cell counter (ThermoFisher).

[0376] Cancer Cell Line Encyclopedia and Connectivity Map Data and Analysis

[0377] Pharmacological profiling data were downloaded from the Cancer Cell Line Encyclopedia (CCLE) data portal (available on the World Wide Web at portals.broadinstitute.org / ccle / data). Data were analyzed and plotted for cell lines derived from lung and breast tissues exposed to the Abl inhibitor AZD0530 and nilotinib. Connectivity map (CMAP) data were downloaded from the CLUE command app within the CMAP data portal (available on the World Wide Web at clue.io / command?q= / home). Top connections and connectivity scores (obtained using the / conn command) were exported between the MEK inhibitor perturbation class (CP_MEK_inhibitor) and the HSP inhibitor perturbation class (CP_HSP_inhibitor) across all cell lines (summary) or across individual cell lines overlapping with this study (A549 and MCF7). Results were then filtered for data from inhibitor exposure. To determine how connectivity varies across all cell lines and across individual cell lines, we filtered for top connections that overlapped with the connectivity summary in the data from individual cell lines. Similar to the related CMAP study (11), we applied a threshold of 90 to the connectivity score.

[0378] Flow cytometry

[0379] A549 and MCF7 cells were seeded in 6 cm dishes at 1.6 x 10 cells per plate. K562 cells were seeded in T25 cm flasks at 1.6 x 10 cells per flask. After 24 h, A549 and MCF7 adherent cells were exposed to 10 μM avinostat, 10 μM pranostin, or DMSO as a vehicle control for 24 h. Treated cells were harvested as described above, washed twice in PBS, resuspended in 500 μL of cold PBS, and fixed by adding 5 mL of ice-cold ethanol while vortexing slowly. Cells were stored at -20°C before processing for flow cytometry analysis. For flow cytometry, the ethanol was removed, washed twice with PBS containing 1% BSA (PBS-B), and blocked for 1 h at room temperature. The blocking buffer was then removed, and the cells were incubated for 2 hours at room temperature in PBS containing 1% BSA and 0.1% Triton X-100 (PBS-BT) and a 1:500 dilution of mouse anti-acetyl-lysine antibody (catalog no. ICP0390, ImmuneChem Pharmaceuticals). After incubation, the cells were washed twice with PBS-BT and incubated with goat anti-mouse Alexa-647 in PBS-BT for 1 hour at room temperature. Finally, the cells were washed twice with PBS-BT and once with PBS-B and resuspended in PBS-B containing 5 μg / ml Hoechst 33258 (Life Sciences Technologies) to stain DNA. The levels of total acetylated lysine and DNA content were then analyzed by flow cytometry on an LSRII flow cytometer (BD Biosciences). Quantification and downstream analysis were performed using FlowJo10 (FlowJo).

[0380] Cell harvesting, nuclei isolation, and sample hashing

[0381] To harvest adherent cells, the medium was removed, the cells were rinsed with 100 μL of DPBS, and Trispin-treated with 50 μL of Tryp-LE for 15 minutes at 37°C. Once the cells detached from the culture plate, the reaction was stopped with 150 μL of ice-cold DMEM containing 10% FBS. A cell suspension was generated by pipetting, and the entire volume was transferred to a 96-well V-bottom plate. The cells were then pelleted by centrifugation at 300 × g for 6 minutes, washed with 100 μL of ice-cold DPBS, and repelleted at 300 × g for 6 minutes.

[0382] Lysis was performed in 96-well V-bottom plates. After removing the PBS, the cell suspension was lysed in 50 μL of cold lysis buffer (10 mM Tris-HCl, pH 7.4, 10 mM NaCl, 3 mM MgCl, 0.1% IGEPAL) supplemented with 1% Superase RNA inhibitor and 400 femtomoles of hash oligo of the form 5'-GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG-[10 bp-barcode]-BAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA-3' (SEQ ID NO: 1), where B is G, C, or T (IDT). The oligos were labeled with CA-630 (24). For screening of larger compounds, 500 femtomoles of additional oligos were used to uniquely index each 96-well treatment plate. After lysis using a multichannel pipette three times, cells were fixed by adding 200 μL of fixation buffer (5% paraformaldehyde, 1.25x PBS). Nuclei were then fixed on ice for 15 min before pooling in a trough. Nuclei were pooled from the plate into a 50 mL conical tube and pelleted by centrifugation at 500 × g for 5 min. Subsequently, cells were resuspended in 500 μL of nuclear suspension buffer (NSB; (10 mM Tris-HCl, pH 7.4, 10 mM NaCl, 3 mM MgCl, 1% Superase 4). RNA inhibitor, 1%, 0.2 mg / mL ultrapure BSA). Finally, nuclei from all plates were pooled into a single conical tube and pelleted by centrifugation at 500 × g for 5 minutes. Nuclei were then resuspended in 1 mL of NSB and flash-frozen in 100 μL aliquots in liquid nitrogen. Nuclei were then stored at -80°C until further processing by sci-RNA-seq.

[0383] sci-RNA-seq2 library preparation

[0384] Frozen nuclei were thawed on ice and spun down at 500 g for 5 min. Next, cells were permeabilized in permeabilization buffer (NSB + 0.25% Triton-X) for 3 min and then spun down. After further washing with NSB, two-level sci-RNA-seq libraries were prepared as previously described (25). Briefly, nuclei were pelleted at 500 × g for 5 min and resuspended in 100 μL of NSB. Cell counts were obtained by staining nuclei with 0.4% trypan blue (Sigma-Aldrich) and counting using a hemocytometer. Next, 5,000 nuclei in 2 μL of NSB and 0.25 μL of 10 mM dNTP mix (Thermo Fisher Scientific, catalog number R0193) were distributed into a skirted twin.tec 96-well LoBind plate (Fisher Scientific, catalog number 0030129512), after which 1 μL of uniquely indexed oligo-dT (25 μM) (25) was added to all wells, incubated at 55°C for 5 min, and placed on ice. Next, 1.75 μL of reverse transcription mix (1 μL of Superscript IV first-strand buffer, 0.25 μL of 100 mM DTT, 0.25 μL of Superscript IV, and 0.25 μL of RNAseOUT recombinant ribonuclease inhibitor) was added to all wells, and the plate was incubated at 55°C for 10 min and placed on ice. To stop the reaction, 5 μL of stop solution (40 mM EDTA, 1 mM spermidine, and 0.5% BSA) was added to each well. Wells were pooled using a wide-bore tip and transferred to a flow cytometry tube through a 0.35 μm filter cap and DAPI added to a final concentration of 3 μM. Pooled nuclei were then sorted into 96-well LoBind plates containing 5 μL of EB buffer (Qiagen) on a FACS Aria II cell sorter (BD) at 150 cells per well. After sorting, 0.75 μL of second-strand mix (0.5 μL mRNA second-strand synthesis buffer and 0.25 μL mRNA second-strand synthesis enzyme, New England Biolabs) was added to each well, and second-strand synthesis was carried out at 16°C for 150 min.Tagging was performed by adding 5.75 μL of tagging mix (0.01 μL of custom TDE1 enzyme, Illumina, in 5.74 μL 2x Nextera TD buffer), and the plate was incubated at 55°C for 5 minutes. The reaction was stopped by adding 12 μL of DNA binding buffer (Zymo) and incubated at room temperature for 5 minutes. DNA was purified using the standard Ampure XP protocol (Beckman Coulter) by adding 36 μL of Ampure XP beads to all wells and eluting with 17 μL of EB buffer. The DNA was then transferred to a new 96-well LoBind plate. For PCR, 2 μL of indexed P5, 2 μL of indexed P7 (25), and 20 μL of NEBNext High-Fidelity master mix were added. A PCR mix (New England Biolabs) was added to each well, and PCR was performed as follows: 75°C for 3 minutes, 98°C for 30 seconds, 18 cycles of 98°C for 10 seconds, 66°C for 30 seconds, and 72°C for 1 minute, followed by a final extension at 72°C for 5 minutes. After PCR, all wells were pooled, concentrated using a DNA cleanup and concentration kit (Zymo), and purified with 0.8X Ampure XP cleanup. Final library concentration was determined by Qubit (Invitrogen), libraries visualized using TapeStation D1000 DNA Screen tape (Agilent), and libraries sequenced on a Nextseq 500 (Illumina) using a high-throughput 75-cycle kit (read 1:18 cycles, read 2:52 cycles, index 1:10 cycles, and index 2:10 cycles).

[0385] Preparation of sci-RNA-seq3 libraries

[0386] Frozen nuclei were thawed as before, and three-level sci-RNA-seq libraries were prepared as described (21). Nuclei were pelleted at 500 × g for 5 min, washed three times with NSB, and a small aliquot of nuclei was stained with 0.4% trypan blue (Sigma-Aldrich). Nuclei were counted using a hemocytometer. Eighty thousand nuclei in 22 μL of NSB, 2 μL of 10 mM dNTP mix, and 2 μL of skirted, ligation-compatible, indexed oligo-dT primer were dispensed into each well of a 96-well LoBind plate, incubated at 55°C for 5 min, and placed on ice. Next, 14 μL of reverse transcription mix (8 μL Superscript IV first-strand buffer, 2 μL 100 mM DTT, 2 μL Superscript IV, and 2 μL RNAseOUT recombinant ribonuclease inhibitor) was added to all wells, and RT was performed in a thermocycler using the following program: 4°C for 2 min, 10°C for 2 min, 20°C for 2 min, 30°C for 2 min, 40°C for 2 min, 50°C for 2 min, and 55°C for 15 min. After RT, 60 μL of nuclear buffer containing BSA (NBB, 10 mM Tris-HCl, pH 7.4, 10 mM NaCl, 3 mM MgCl, and 1% BSA) was added to each well, and the nuclei were pooled using a wide-bore tip and centrifuged at 500 × g for 10 min to pellet the nuclei. The supernatant was removed. A second round of combinatorial indexing was performed by ligating an indexed primer to the 5' end of the RT-indexed cDNA. Nuclei were resuspended in NSB and 10 μL was added to each well of a 96-well LoBind plate. Then, 8 μL of indexed ligation primer was added to each well along with 22 μL of ligation mix (20 μL of Quick ligase buffer and 2 μL of Quick ligase, New England Biolabs). Ligation was then carried out at 25°C for 10 minutes. After ligation, 60 μL of NBB was added to each well, and the nuclei were pooled using a wide-bore tip. An additional 40 mL of NBB was added to the nuclei and centrifuged at 600 × g for 10 minutes to pellet the nuclei. The supernatant was removed.Nuclei were then washed once with 5 mL of NBB, resuspended in 4 mL of NBB, and filtered through a 40 μm Flowmi cell strainer (Sigma-Aldrich) to remove multiplets. Nuclei were counted and distributed into 96-well LoBind plates in a volume of 5 μL, 5,000 nuclei per well. The plates containing the nuclei were frozen and stored at -80°C until further processing. After thawing the frozen plates, 5 μL of second-strand synthesis mix (3 μL elution buffer, 1.33 μL mRNA second-strand synthesis buffer, and 0.66 μL mRNA second-strand synthesis enzyme) was added to each well and incubated at 16°C for 3 hours. Tagging was performed by adding 10 μL of tagging mix (0.01 μL custom TDE1 enzyme in 9.99 μL 2x Nextera TD buffer, Illumina), and the plate was incubated at 55°C for 5 minutes. After ta...

Claims

1. An oligo composition for normalizing a sequencing library, The oligo composition comprises a plurality of oligos, The oligo composition comprises a certain number of subgroups of the plurality of oligos, the number of subgroups being in the range of 2 to 100, and the oligos within the subgroups include PCR handles or universal sequences. Each subgroup of the oligo composition includes an oligo containing an index sequence that is unique to the oligo of that subgroup and does not exist on the oligo of one or more other subgroups of the oligo composition. An oligo composition wherein at least two subpopulations of oligos in the oligo composition are measurably different from each other and have a number of oligos that provide a concentration ladder when cells or nuclei are normalized in the sequencing library.

2. The oligo composition according to claim 1, wherein the number of subgroups is in the range of 2 to 16.

3. The oligo composition according to claim 1, wherein at least two subgroups of oligos in the oligo composition have the same number of oligos.

4. The oligo composition according to claim 1, wherein all of the subgroups of oligos in the oligo composition have a number of oligos that are measurably different from one another.

5. The oligo composition according to any one of claims 1 to 4, wherein the oligo composition comprises a group of 2 to 8 oligos.

6. The oligo composition according to claim 1, wherein the oligo comprises single-stranded DNA.

7. The oligo composition according to claim 1, wherein the oligo includes a unique molecular identifier.

8. The oligo composition according to claim 1, wherein the oligo comprises a universal sequence.

9. The oligo composition according to claim 1, wherein the oligo comprises a non-nucleic acid component.

10. The oligo composition according to claim 9, wherein the non-nucleic acid component comprises a protein.

11. The oligo composition according to claim 1, wherein at least two subgroups of oligos include a first subgroup of oligos and a second subgroup of oligos, and the number of oligos in the second subgroup is one order of magnitude greater than the number of oligos in the first subgroup.

12. The oligo composition according to claim 1, wherein at least two subgroups of oligos include a first subgroup of oligos and a second subgroup of oligos, and the number of oligos in the second subgroup is greater than the number of oligos in the first subgroup by a coefficient of more than 1 and up to 10,000.

13. A plurality of compartments, each compartment containing the oligo composition described in Claim 1.

14. The plurality of compartments according to claim 13, wherein each compartment includes a well or a droplet.

15. The plurality of compartments according to claim 13, wherein each compartment further comprises a nucleus or cell, and the at least two subpopulations of oligos for normalizing the sequencing library associate with the nucleus or cell.

16. The plurality of sections according to claim 13, wherein the number of oligos in each subgroup is in the range of at least 0.001 zeptomoles to 100 atomoles.

17. A population of nuclei or cells, wherein the nuclei or cells comprise the oligo composition described in Claim 1, and each subpopulation of oligos for normalizing the sequencing library associates with the nuclei or cells.

18. The population according to claim 17, wherein the association between the nucleus or cell and the oligo for normalizing the sequencing library is nonspecific.

19. A method for normalizing a sequencing library containing nucleic acids from a plurality of single nuclei or single cells, wherein the method is: (a) Providing a first plurality of compartments containing isolated nuclei or cells, (b) A step of contacting the isolated nuclei or cells of each compartment with the oligo composition according to claim 1, wherein members of each subpopulation of oligos for normalizing the sequencing library are associated with the isolated nuclei or cells, (c) A method comprising the step of combining labeled nuclei or labeled cells from different compartments to generate pooled labeled nuclei or pooled labeled cells.

20. The method according to claim 19, further comprising the steps of exposing the cells in each compartment to predetermined conditions, or exposing the cells in each compartment to predetermined conditions and then isolating nuclei from a plurality of cells prior to step (a).

21. The method according to claim 20, wherein the predetermined conditions include exposure to a drug.

22. The method according to claim 21, wherein the agent comprises a protein, a non-ribosomal protein, a polyketide, an organic molecule, an inorganic molecule, RNA or RNAi molecule, a carbohydrate, a glycoprotein, a nucleic acid, a drug, or a combination thereof.

23. A step prior to step (b), wherein the isolated nuclei or cells of each compartment are brought into contact with the hashing oligo, At least one copy of the hashing oligo associates with an isolated nucleus or cell. The hashing oligo comprises nucleic acid and hashing index, The hashing index within each compartment includes an index sequence that is different from the index sequences in other compartments and different from the index sequences of the oligos for normalizing the sequencing library present in the compartment, in order to generate labeled hash nuclei or labeled hash cells, the step of The method according to claim 19, further comprising the step of combining the labeled hash nuclei or labeled hash cells from different compartments to generate a pooled labeled hash nuclei or pooled labeled hash cells.

24. The method of claim 19, further comprising the step of exposing the cells or nuclei to a crosslinking compound in order to fix the oligos for normalizing the sequencing library to the cells or isolated nuclei.

25. The method according to claim 24, wherein the crosslinking compound comprises paraformaldehyde, formalin, or methanol.

26. The method according to claim 19, wherein the association between the oligo for normalizing the sequencing library and the cell or isolated nucleus is nonspecific.

27. ​​The method according to claim 26, wherein the nonspecific association between the oligo and the cell or isolated nucleus for normalizing the sequencing library is by absorption.

28. The method according to claim 23, further comprising the step of processing the pooled labeled hash cells or pooled labeled hash nuclei using a single-cell combinatorial indexing method to obtain a sequencing library comprising nucleic acids from the plurality of single nuclei or single cells, wherein the nucleic acids comprise a plurality of indices.

29. The method according to claim 28, wherein the single-cell combinatorial indexing method is single-nuclear transcriptome sequencing, single-cell transcriptome sequencing, single-cell transcriptome and transposon-accessible chromatin sequencing, single-nuclear whole-genome sequencing, single-nuclear sequencing of transposon-accessible chromatin, sci-HiC, drug-seq, sci-CAR, sci-MET, sci-Crop, sci-perturb, or sci-Crispr.

30. A method for normalizing a sequencing library containing nucleic acids from a plurality of single nuclei or single cells, wherein the method is: (a) the step of providing an isolated nucleus or cell, (b) A step of contacting the isolated nuclei or cells with the oligo composition according to claim 1, wherein each subpopulation member of the oligo for normalizing the sequencing library associates with the isolated nuclei or cells to form labeled nuclei or labeled cells, (c) A method comprising the step of distributing a subset of labeled nuclei or labeled cells into a plurality of compartments.

31. The method according to claim 30, further comprising the steps of exposing the cells to predetermined conditions, or exposing the cells to predetermined conditions and then isolating nuclei from a plurality of cells prior to step (a).

32. The method according to claim 31, wherein the predetermined conditions include exposure to a drug.

33. The method according to claim 32, wherein the agent comprises a protein, a non-ribosomal protein, a polyketide, an organic molecule, an inorganic molecule, RNA or RNAi molecule, a carbohydrate, a glycoprotein, a nucleic acid, a drug, or a combination thereof.

34. A step after step (c) of bringing the isolated nuclei or cells of each compartment into contact with the hashing oligo, At least one copy of the hashing oligo associates with an isolated nucleus or cell. The hashing oligo comprises nucleic acid and hashing index, The hashing index within each compartment includes an index sequence that is different from the index sequences in other compartments and different from the index sequences of the oligos for normalizing the sequencing library present in the compartment, in order to generate labeled hash nuclei or labeled hash cells, the step of The method according to claim 30, further comprising the step of combining the labeled hash nuclei or labeled hash cells from different compartments to generate a pooled labeled hash nuclei or pooled labeled hash cells.

35. The method of claim 30, further comprising the step of exposing the cells or nuclei to a crosslinking compound in order to fix the oligos for normalizing the sequencing library to the cells or isolated nuclei.

36. The method according to claim 35, wherein the crosslinking compound comprises paraformaldehyde, formalin, or methanol.

37. The method according to claim 30, wherein the association between the oligo and the cell or isolated nucleus for normalizing the sequencing library is nonspecific.

38. The method according to claim 37, wherein the nonspecific association between the oligo and the cells or isolated nuclei for normalizing the sequencing library is by absorption.

39. The method according to claim 34, further comprising the step of processing the pooled labeled hash cells or pooled labeled hash nuclei using a single-cell combinatorial indexing method to obtain a sequencing library comprising nucleic acids from the plurality of single nuclei or single cells, wherein the nucleic acids comprise a plurality of indices.

40. The method according to claim 39, wherein the single-cell combinatorial indexing method is single-nuclear transcriptome sequencing, single-cell transcriptome sequencing, single-cell transcriptome and transposon-accessible chromatin sequencing, single-nuclear whole-genome sequencing, single-nuclear sequencing of transposon-accessible chromatin, sci-HiC, drug-seq, sci-CAR, sci-MET, sci-Crop, sci-perturb, or sci-Crispr.