High-throughput single-cell libraries and methods of making and of using
The use of a transpososome complex for single-cell combinatorial indexing addresses the challenge of sequencing rare cells by generating comprehensive sequence data without custom transposons, facilitating efficient characterization and enrichment of biological features.
Patent Information
- Application Number
- JP2025069356
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-12-19
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-23
AI Technical Summary
Current single-cell genomic technologies face challenges in comprehensively sequencing rare cells within a population without enrichment, which is costly and difficult, and often require custom-modified transposons for labeling nucleic acid content.
A method using a transpososome complex for single-cell combinatorial indexing that does not necessitate custom-modified transposons, involving the incorporation of universal sequences into DNA nucleic acids, followed by partitioning and indexing steps to generate sequencing libraries.
Enables efficient and cost-effective characterization of rare cell populations by generating comprehensive sequence data, allowing for the identification and enrichment of specific biological features in single-cell sequencing libraries.
Smart Images

Figure 2025108660000009 
Figure 2025108660000010 
Figure 2025108660000011
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications)
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 950,670, filed on Dec. 19, 2019, which is hereby incorporated by reference in its entirety.
[0003] (Government Sponsorship)
[0004] This invention was made with government support under grant number T32 HL007828 awarded by the National Institutes of Health. The government has certain rights in this invention.
[0005] (Field of the Invention)
[0006] Embodiments of the present disclosure relate to nucleic acid sequencing. Specifically, embodiments of the methods and compositions provided herein relate to generating single - cell combinatorial indexed sequencing libraries and then obtaining sequence data. In some embodiments, the sequence data obtained from the library is comprehensive, and in other embodiments, the sequence data obtained from the library enables the characterization of rare events.
[0007]
Background Art
[0008] Single - cell combinatorial indexing (“sci - ”) is a methodological framework that uses split - pool barcoding to uniquely label the nucleic acid content of a large number of single cells or single nuclei to create single - cell combinatorial sequencing libraries. Current single - cell genomic technologies often involve adding unique labels in one step using transpososome complexes, which requires large amounts of custom - modified transposons.
[0009] Single cell genomics techniques resolve cell-to-cell differences that are difficult to determine when studying bulk populations of cells. In many important applications such as oncology, immunology, and metagenomics, there is great interest in characterizing rare cells, and there are challenges. Current methods of single cell sequencing can characterize millions of single cells in parallel. However, comprehensively sequencing-based characterization of rare cells within a population without enrichment is costly and difficult.
[0010] SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0011] Provided herein is a method of using a transpososome complex during single cell combinatorial indexing without the need to generate custom-modified transposons.
[0012] In one embodiment, the present disclosure provides a method for preparing a sequencing library comprising nucleic acids from a plurality of single nuclei or single cells. The method includes providing a plurality of nuclei or cells, wherein the nuclei or cells comprise nucleosomes, and contacting the plurality of nuclei or cells with a transpososome complex comprising a transposase and a universal sequence. In one embodiment, the plurality of nuclei or cells are in bulk upon contact with the transpososome complex, and in another embodiment, upon contact with the transpososome complex, the plurality of nuclei or cells are partitioned within a first plurality of compartments, each compartment comprising a subset of the nuclei or cells or representing a sample. Contacting further includes conditions suitable for incorporating the universal sequence into DNA nucleic acids to yield double-stranded DNA nucleic acids comprising the universal sequence. In embodiments where contacting occurs with the plurality of nuclei or cells in bulk, the method also includes partitioning the plurality of nuclei or cells into a first plurality of compartments, each compartment comprising a subset of the nuclei or cells. The DNA molecules within each subset of the nuclei or cells are processed to generate indexed nuclei or cells. This processing includes adding a first compartment-specific index sequence to the DNA nucleic acids present in each subset of the nuclei or cells to yield indexed nucleic acids present in the indexed nuclei or cells. This processing may include ligation, primer extension, hybridization, amplification, or combinations thereof. The indexed nuclei or cells can be combined to generate pooled indexed nuclei or cells.
[0013] In one embodiment, providing can include providing the plurality of nuclei or cells in a plurality of compartments, each compartment comprising a subset of the nuclei or cells or representing a sample. Contacting can include contacting each compartment with the transpososome complex, and the method can further include combining the nuclei or cells after contact to generate pooled nuclei or cells.
[0014] In one embodiment, contacting comprises contacting each subset with two transpososome complexes, one transpososome complex comprising a first transposase comprising a first universal array, and the second transpososome complex comprising a second transposase comprising a second universal array, and contacting further comprises conditions suitable for incorporating the first universal array and the second universal array into a DNA nucleic acid to yield a double-stranded DNA nucleic acid comprising the first universal array and the second universal array.
[0015] In one embodiment, the method may further comprise distributing a pooled indexed nuclei or cells comprising indexed nuclei or cells into a second plurality of compartments, each compartment comprising a subset of the nuclei or cells, and processing the DNA molecules within each subset of the nuclei or cells to generate doubly-indexed nuclei or cells. Processing may comprise adding a second compartment-specific index sequence to the DNA nucleic acids present in each subset of the nuclei or cells to yield a doubly-indexed nucleic acid present in the indexed nuclei or cells. The method may comprise combining the doubly-indexed nuclei or cells to generate a pooled doubly-indexed nuclei or cells.
[0016] In one embodiment, the method may further comprise distributing a pooled indexed nuclei or cells comprising doubly-indexed nuclei or cells into a third plurality of compartments, each compartment comprising a subset of the nuclei or cells, and processing the DNA molecules within each subset of the nuclei or cells to generate triply-indexed nuclei or cells. Processing may comprise adding a third compartment-specific index sequence to the DNA nucleic acids present in each subset of the nuclei or cells to yield a triply-indexed nucleic acid present in the indexed nuclei or cells. The method may comprise combining the triply-indexed nuclei or cells to generate a pooled triply-indexed nuclei or cells.
[0017] In one embodiment, the method comprises obtaining indexed nucleic acids (e.g., dual-indexed, triple-indexed, etc.) from pooled indexed nuclei or cells, and may further include generating a sequencing library from a plurality of nuclei or cells.
[0018] Also provided herein are methods for identifying and / or characterizing subpopulations of cells. In one embodiment, the method includes providing a sequencing library, such as a single-cell combinatorial sequencing library. Optionally, the sequencing library is generated from a population of cells or nuclei enriched for a characteristic. The method may include interrogating the sequencing library by targeted sequencing. Targeted sequencing may be based on a biological feature typically present in a small percentage of the cells used to generate the library. Examples of biological features include, but are not limited to, cell class, species type, or nucleotide sequences indicative of a disease state. In addition to targeted sequencing of the biological feature, the sequencing also includes determining the sequence of an index sequence present in the same modified target nucleic acid as the biological feature. As a result, members of the sequencing library derived from the same cells or nuclei as members of the library containing the biological feature are identified. The method further includes modifying the sequencing library to increase the representation of these members derived from the same cells or nuclei as members of the library containing the biological feature. This modification may include enriching for desired members of the sequencing library or depleting undesired members of the sequencing library to yield a sublibrary.
[0019] Definitions
[0020] The terms used herein will be understood to have their ordinary meaning in the relevant art, unless otherwise specified. Some of the terms used herein and their meanings are set forth below.
[0021] As used herein, the terms "organism" and "subject" are used interchangeably and refer to microorganisms (e.g., prokaryotes or eukaryotes), animals, and plants. Examples of animals are mammals such as humans.
[0022] As used herein, the term "cell type" is intended to identify cells based on morphology, phenotype, origin of development, or other known or recognizable distinguishable cell characteristics. A variety of different cell types can be obtained from a single organism (or organisms of the same species). Exemplary cell types include gametes (including, for example, female gametes such as ova or egg cells, and male gametes such as sperm), ovarian epithelium, ovarian epithelium, ovarian fibroblasts, testis, bladder, immune cells, B cells, T cells, natural killer cells, dendritic cells, cancer cells, eukaryotic cells, stem cells, blood cells, muscle cells, adipocytes, skin cells, nerve cells, bone cells, pancreatic cells, endothelial cells, pancreatic β, pancreatic endothelium, bone marrow lymphoblasts, bone marrow lymphoblasts, bone marrow macrophages, myeloblasts, bone marrow adipocytes, bone marrow osteoblasts, bone marrow chondrocytes, bone marrow chondrocytes, bone marrow chondrocytes, myeloblasts, bone marrow chondrocytes, promyeloblasts, megakaryoblasts of bone marrow, bladder, brain B lymphocytes, brain glial cells, neurons, brain astrocytes, neuroectoderm, brain macrophages, microglia of brain, brain epithelium, cortical neurons, brain fibroblasts, breast epithelium, colonic epithelium, colonic B lymphocytes, mammary epithelium, mammary myoepithelium, mammary fibroblasts, colonic enterocytes, cervical epithelium, ductal epithelium of breast, tongue epithelium, dendritic cells of tonsil, tonsil lymphocytes, peripheral blood lymphoblasts, peripheral blood T lymphoblasts, peripheral blood T lymphoblasts, peripheral blood natural killer, peripheral blood B lymphoblasts, peripheral blood monocytes, peripheral blood myeloblasts, peripheral blood monoblasts, peripheral blood monoblasts, peripheral blood monoblasts, peripheral blood monoblasts, peripheral blood T lymphocytes, peripheral blood promyeloblasts, peripheral blood macrophages, peripheral blood basophils, liver endothelium, liver mast, liver epithelium, liver B lymphocytes, spleen endothelium, spleen epithelium, spleen B lymphocytes, hepatocytes, liver, fibroblasts, lung epithelium, bronchial epithelium, lung fibroblasts, lung B lymphocytes, lung Schwann, lung squamous epithelium, lung macrophages, lung osteoblasts, neuroendocrine, pulmonary alveoli, gastric epithelium, and gastric fibroblasts, but are not limited thereto. In one embodiment, the various different cell types obtained from a single organism can include other cells such as the cells of the organism and the cells of symbiotic or pathogenic microorganisms associated with the organism. Examples of symbiotic or pathogenic microorganisms associated with an organism include, but are not limited to, prokaryotic and eukaryotic microorganisms present in a microbiome sample derived from the organism or present within a tissue, optionally causing disease.
[0023] As used herein, the term "tissue" is intended to mean a collection or aggregate of cells that act together to perform one or more specific functions within an organism. The cells may optionally be morphologically similar. Exemplary tissues include, but are not limited to, embryonic, pituitary, eye, muscle, skin, tendon, vein, artery, blood, heart, spleen, lymph node, bone, bone marrow, lung, bronchus, trachea, intestine, small intestine, large intestine, colon, rectum, salivary gland, tongue, gallbladder, appendix, liver, pancreas, brain, stomach, skin, kidney, ureter, bladder, urethra, gonad, testis, ovary, uterus, fallopian tube, thymus, pituitary gland, thyroid gland, adrenal gland, or parathyroid gland. Tissues can be derived from any of a variety of organs of a human or other organism. Tissues can be healthy or unhealthy. Examples of unhealthy tissues include malignant tumors of reproductive tissue, lung, breast, colorectal, prostate, nasopharynx, stomach, testis, skin, nervous system, bone, ovary, liver, blood tissue, pancreas, uterus, kidney, lymphoid tissue, etc. Malignant tumors can be of various histological subtypes, such as carcinomas, adenocarcinomas, sarcomas, fibroadenocarcinomas, neuroendocrine, or undifferentiated, but are not limited thereto.
[0024] As defined herein, "sample" and its derivatives are used in their broadest sense and include any specimen, culture, etc. suspected of containing a target nucleic acid and / or a target protein. In some embodiments, the sample comprises DNA, RNA, protein, or a combination thereof. The sample can include any biological sample, clinical sample, surgical sample, agricultural sample, atmospheric sample, or aquatic-based sample containing one or more nucleic acids and / or one or more proteins. The term also includes any isolated nucleic acid from a sample, such as genomic DNA or transcriptome, and any isolated protein from a sample. In some embodiments, the sample includes a collection of cells or nuclei.
[0025] As used herein, the term "compartment" is intended to mean an area or volume that separates or isolates something from other things. Exemplary compartments include, but are not limited to, vials, tubes, wells, droplets, boluses, beads, containers, surface features, or areas or volumes separated by physical forces such as fluid flow, magnetism, electric current. In one embodiment, the compartment is a well of a multi-well plate such as a 96- or 384-well plate. In one embodiment, the compartment is a well of a patterned surface (e.g., a microwell or nanowell). As used herein, a droplet can be a bead for encapsulating one or more nuclei or cells and can include hydrogel beads containing a hydrogel composition. In some embodiments, the droplet is a homogeneous droplet of a hydrogel material or a hollow droplet having a polymeric hydrogel shell. Regardless of whether it is homogeneous or hollow, the droplet can be capable of encapsulating one or more nuclei or cells. In some embodiments, the droplet is a surfactant-stabilized droplet.
[0026] As used herein, "transpososome complex" refers to a nucleic acid comprising an integrase and an integration recognition site. A "transpososome complex" is a functional complex formed by a transposase and a transposase recognition site capable of catalyzing a transposition reaction (see, e.g., Gunderson et al., WO 2016 / 130704). Examples of integrases include, but are not limited to, integrases or transposases. Examples of integration recognition sites include, but are not limited to, transposase recognition sites.
[0027] As used herein, the term "nucleic acid" is used interchangeably with polynucleotide and oligonucleotide. Nucleic acids are intended to be consistent with their use in the art and include naturally occurring nucleic acids or functional analogs thereof. Particularly useful functional analogs are those that can hybridize to nucleic acids in a sequence-specific manner or can be used as a template for replicating a specific nucleotide sequence. Naturally occurring nucleic acids generally have a backbone containing phosphodiester bonds. Analogous structures can have alternative backbone linkages, including any of a variety of those known in the art. Naturally occurring nucleic acids generally have deoxyribose sugars (e.g., found in deoxyribonucleic acid (DNA)) or ribose sugars (e.g., found in ribonucleic acid (RNA)). Nucleic acids can contain any of a variety of analogs of these sugar moieties known in the art. Nucleic acids can contain natural or non-natural bases. In this regard, naturally occurring deoxyribonucleic acid can have one or more bases selected from the group consisting of adenine, thymine, cytosine, or guanine, and ribonucleic acid can have one or more bases selected from the group consisting of adenine, uracil, cytosine, or guanine. Useful non-natural bases that can be included in nucleic acids are known in the art. Examples of non-natural bases include locked nucleic acids (LNA), bridged nucleic acids (BNA), and pseudo-complementary bases (Trilink Biotechnologies, San Diego, California). Incorporating LNA and BNA bases into DNA oligonucleotides can enhance the hybridization strength and specificity of the oligonucleotides. LNA and BNA bases, and the use of such bases, are known and routine to those of skill in the art. Unless otherwise stated, the term "nucleic acid" includes natural and non-natural DNA, mRNA, and non-coding RNA, such as RNA without a polyA at the 3' end, and nucleic acids derived from RNA, such as cDNA. The term "nucleic acid" refers only to the primary structure of the molecule. Thus, the term includes triple-stranded, double-stranded, and single-stranded deoxyribonucleic acid ("DNA"), and triple-stranded, double-stranded, and single-stranded ribonucleic acid ("RNA").
[0028] As used herein, the term "target" is intended as a semantic identifier of a molecule whose source, function, identity, and / or composition is being investigated. Examples of targets include, but are not limited to, nucleic acids and proteins. As used herein, when the term "target" is used with respect to a nucleic acid, it is intended as a semantic identifier of the nucleic acid in the context of the methods or compositions described herein, and does not necessarily limit the structure or function of the nucleic acid other than as explicitly indicated otherwise. A target nucleic acid may be any nucleic acid of essentially known or unknown sequence. This may be, for example, a fragment of genomic DNA (e.g., chromosomal DNA), extrachromosomal DNA such as a plasmid, cell-free DNA, RNA (e.g., mRNA or non-coding RNA), protein (e.g., cellular or cell surface protein), or cDNA. A target nucleic acid may be a nucleic acid that binds to a compound such as an antibody that specifically binds to a biomolecule such as a protein, glycan, proteoglycan, or lipid (U.S. Patent Application Publication No. 2018 / 0273933). Sequencing can result in the determination of the sequence of all or part of the target molecule. The target may be derived from a primary nucleic acid sample such as a nucleus. In one embodiment, the target can be processed into a template suitable for amplification by placing universal sequences at one or both ends of each target fragment. The target can also be obtained from a primary RNA sample by reverse transcription into cDNA. In one embodiment, the target is used with reference to a subset of DNA, RNA, or protein present within a cell. Target sequencing typically uses selection and isolation of the gene or region or protein of interest by either PCR amplification (e.g., region-specific primers) or a hybridization-based capture method or an antibody. Target enrichment can be performed at various stages of the method. For example, target RNA expression can be obtained by using target-specific primers in the reverse transcription step or by hybridizing-based enrichment of a subset from a more complex library. Examples include exome sequencing or the L1000 assay (Subramanian et al., 2017, Cell, 171; 1437-1452).Target sequencing can include any of the enrichment processes known to those of ordinary skill in the art. A target nucleic acid having one or both ends of a universal array may be referred to as a modified target nucleic acid. References to nucleic acids such as target nucleic acids include both single-stranded and double-stranded nucleic acids unless otherwise specified. In one embodiment, the library is enriched using an index sequence or a plurality of index sequences. In some embodiments, the enrichment includes one or more index sequences attached to the same library molecule and is introduced, for example, via combinatorial indexing.
[0029] As used herein, the term "universal," when used to describe a nucleotide sequence, refers to a region of a sequence that is common to two or more nucleic acid molecules, and the molecules also have regions of sequences that are different from each other. Universal sequences present in different members of a collection of molecules, such as members of a sequencing library, can enable the capture of multiple different nucleic acids using a population of universal capture sequences. Non-limiting examples of universal capture sequences include sequences that are identical or complementary to the P5 and P7 primers. Similarly, universal sequences present in different members of a collection of molecules can be used to replicate (e.g., sequence) or amplify multiple different nucleic acids using a population of universal primers that are complementary to a portion of the universal sequence, such as a universal primer binding site. The terms "A14" and "B15" can be used when referring to a universal primer binding site. The terms "A14'" (A14 prime) and "B15'" (B15 prime) refer to the complements of A14 and B15, respectively. It will be understood that any suitable universal primer binding site can be used in the methods presented herein, and the use of A14 and B15 is merely exemplary embodiments. In one embodiment, the universal primer binding site is used as the site where a universal primer (e.g., a sequencing primer for Read 1 or Read 2) anneals for sequencing.
[0030] The terms "P5" and "P7" can be used when referring to a universal capture array or capture oligonucleotide. The terms "P5'" (P5 prime) and "P7'" (P7 prime) refer to the complements of P5 and P7, respectively. In the methods presented herein, any suitable universal capture array or capture nucleotide can be used, and it will be understood that the use of P5 and P7 is only an exemplary embodiment. The use of capture nucleotides such as P5 and P7 or their complements on a flow cell is known in the art as exemplified by the disclosures of International Publication Nos. WO 2007 / 010251, WO 2006 / 064199, WO 2005 / 065814, WO 2015 / 106941, WO 1998 / 044151, and WO 2000 / 018957. For example, any suitable forward amplification primer can be useful in the methods presented herein for complementary sequences and amplification of sequences, whether immobilized or in solution. Similarly, any suitable reverse amplification primer can be useful in the methods presented herein for complementary sequences and amplification of sequences, whether immobilized or in solution. One of ordinary skill in the art will understand how to design and use primer sequences suitable for the capture and / or amplification of nucleic acids presented herein.
[0031] As used herein, the term "primer" and its derivatives generally refer to any nucleic acid capable of hybridizing to a target sequence. Typically, a primer functions as a substrate to which nucleotides can be polymerized by a polymerase or to which nucleotide sequences such as indexes can be ligated. In some embodiments, however, a primer can be incorporated into a synthesized nucleic acid strand and provide a site where another primer can hybridize to prime the synthesis of a new strand complementary to the synthesized nucleic acid molecule. A primer can comprise any combination of nucleotides or their analogs. A primer can be single-stranded, double-stranded, or a nucleic acid containing single-stranded and double-stranded regions, and can include ribonucleotides, deoxyribonucleotides, their analogs, or mixtures thereof. The terms "polynucleotide" and "oligonucleotide" are used interchangeably herein. These terms include, as equivalents, any analog of DNA, RNA, cDNA, or antibody-oligonucleotide complex made from nucleotide analogs, and are understood to be applicable to single-stranded (such as sense or antisense) and double-stranded polynucleotides. This term as used herein also encompasses cDNA, which is complementary or copy DNA produced from an RNA template, for example by the action of reverse transcriptase. This term refers only to the primary structure of the molecule. Thus, this term includes triple-stranded, double-stranded, and single-stranded deoxyribonucleic acid ("DNA"), as well as triple-stranded, double-stranded, and single-stranded ribonucleic acid ("RNA").
[0032] As used herein, the term "adapter" and its derivatives, such as "universal adapter", generally refer to any linear oligonucleotide that can bind to the nucleic acid molecules of the present disclosure. In some embodiments, the adapter is substantially non-complementary to the 3' or 5' end of any target sequence present in the sample. In some embodiments, a suitable adapter length ranges from about 10 - 100 nucleotides, from about 12 - 60 nucleotides, or from about 15 - 50 nucleotides in length. Generally, an adapter can comprise any combination of nucleotides and / or nucleic acids. In some aspects, the adapter can comprise one or more cleavable groups at one or more positions. In another aspect, the adapter can comprise a sequence that is substantially identical to or substantially complementary to at least a portion of a primer, such as a universal primer. In some embodiments, the adapter can comprise a barcode (also referred to herein as a tag or index) to assist with downstream error correction, identification, or sequencing. The terms "adaptor" and "adapter" are used interchangeably.
[0033] As used herein, the term "each", when used with respect to a set of items, is intended to identify the individual items within the set, but does not necessarily refer to all of the items within the set, unless the context clearly indicates otherwise.
[0034] As used herein, the term "transport" refers to the movement of molecules through a fluid. This term can include passive transport, such as the movement of molecules along their concentration gradient (e.g., passive diffusion). This term can also include active transport, where molecules can move along or against their concentration gradient. Thus, transport can include applying energy to move one or more molecules in a desired direction or to a desired location, such as an amplification site.
[0035] As used herein, "amplification," "amplify," or "amplification reaction" and their derivatives generally refer to any action or process in which at least a portion of a nucleic acid molecule is replicated or copied into at least one additional nucleic acid molecule. The additional nucleic acid molecule optionally contains a sequence that is substantially identical to or substantially complementary to at least a portion of the template nucleic acid molecule. The template nucleic acid molecule may be single-stranded or double-stranded, and the additional nucleic acid molecule may independently be single-stranded or double-stranded. Amplification optionally includes linear or exponential replication of the nucleic acid molecule. In some embodiments, such amplification can be performed using isothermal conditions, and in other embodiments, such amplification may include thermal cycling. In some embodiments, amplification is multiplex amplification that includes simultaneous amplification of multiple target sequences in a single amplification reaction. In some embodiments, "amplification" includes amplifying at least a portion of DNA- and RNA-based nucleic acids, either alone or in combination. The amplification reaction may include any of the amplification processes known to those of skill in the art. In some embodiments, the amplification reaction includes polymerase chain reaction (PCR).
[0036] As used herein, "amplification conditions" and derivatives thereof generally refer to conditions suitable for amplifying one or more nucleic acid sequences. Such amplification can be linear or exponential. In some embodiments, the amplification conditions can include isothermal conditions, or can include thermal cycling conditions, or a combination of isothermal and thermal cycling conditions. In some embodiments, conditions suitable for amplifying one or more nucleic acid sequences include polymerase chain reaction (PCR) conditions. Typically, the amplification conditions refer to a reaction mixture sufficient to amplify a nucleic acid such as one or more target sequences flanked by universal sequences, or a reaction mixture sufficient to amplify an amplified target sequence ligated to one or more adapters. Generally, the amplification conditions include a catalyst for amplification, or nucleic acid synthesis, such as a polymerase, a primer having some complementarity to the nucleic acid to be amplified, and nucleotides such as deoxyribonucleotide triphosphates (dNTPs) to promote extension of the primer when hybridized to the nucleic acid. The amplification conditions may require hybridization or annealing of the primer to the nucleic acid, extension of the primer, and a denaturation step in which the extended primer is separated from the nucleic acid sequence undergoing amplification. Typically, but not necessarily, the amplification conditions may include thermal cycling, but in some embodiments, the amplification conditions include a plurality of cycles in which the steps of annealing, extension, and separation are repeated. Typically, the amplification conditions include cations such as Mg 2+ or Mn 2+ and may also include various modifiers of ionic strength.
[0037] As used herein, "re-amplification" and derivatives thereof generally refer to any process (referred to in some embodiments as "secondary" amplification) in which at least a portion of an amplified nucleic acid molecule is further amplified via any suitable amplification process, thereby generating a re-amplified nucleic acid molecule. The secondary amplification need not be the same as the original amplification process by which the amplified nucleic acid molecule was generated, and the re-amplified nucleic acid molecule need not be completely identical or completely complementary to the amplified nucleic acid molecule; all that is required is that the re-amplified nucleic acid molecule contain at least a portion of the amplified nucleic acid molecule or its complement. For example, re-amplification can involve the use of different amplification conditions and / or different primers, including target-specific primers that are different from the primary amplification.
[0038] As used herein, the term "polymerase chain reaction" ("PCR") refers to the method of Mullis (U.S. Patent Nos. 4,683,195 and 4,683,202) that describes a method for increasing the concentration of a segment of a polynucleotide of interest in a mixture of genomic DNA without cloning or purification. This process for amplifying a polynucleotide of interest consists of introducing a large excess of two oligonucleotide primers into a DNA mixture containing the desired polynucleotide of interest, followed by a series of thermal cycling in the presence of a DNA polymerase. The two primers are complementary to each of the strands of the target double-stranded polynucleotide. First, the mixture is denatured at a higher temperature, and then the primers are annealed to complementary sequences within the polynucleotide of the molecule of interest. After annealing, the primers are extended with polymerase to form new pairs of complementary strands. The steps of denaturation, primer annealing, and polymerase extension can be repeated many times (referred to as thermal cycling) to obtain a high concentration of amplified segments of the desired polynucleotide of interest. The length of the amplified segment (amplicon) of the desired polynucleotide of interest is determined by the relative positions of the primers with respect to each other, and thus this length is a controllable parameter. By repeating this process, this method is called PCR. The desired amplified segments of the polynucleotide of interest become the major nucleic acid sequences (with respect to concentration) in the mixture and are thus said to be "PCR amplified." In a modification of the above method, target nucleic acid molecules can be PCR amplified using multiple different primer pairs, and in some cases, one or more primer pairs per target nucleic acid molecule of interest can be used for PCR amplification, thereby forming a multiplex PCR reaction.
[0039] As defined herein, "multiplex amplification" refers to the selective and non-random amplification of two or more target sequences in a sample using at least one target-specific primer. In some embodiments, multiplex amplification is performed such that some or all of the target sequences are amplified within a single reaction vessel. The "plex" of a given multiplex amplification generally refers to the number of different target-specific sequences that are amplified during that single multiplex amplification. In some embodiments, the plex can be about 12 plex, 24 plex, 48 plex, 96 plex, 192 plex, 384 plex, 768 plex, 1536 plex, 3072 plex, 6144 plex, or more. The amplified target sequences can also be detected by several different methodologies (e.g., gel electrophoresis followed by densitometry, quantification by a bioanalyzer or quantitative PCR, hybridization with a labeled probe, incorporation of a biotinylated primer followed by detection of an avidin-enzyme conjugate, incorporation of 32 P-labeled deoxynucleotide triphosphates).
[0040] As used herein, "amplified target sequence" and its derivatives generally refer to a polynucleotide sequence produced by amplifying a target sequence using a target-specific primer and the methods provided herein. The amplified target sequence can be either in the same sense (i.e., the plus strand) or the antisense (i.e., the minus strand) with respect to the target sequence.
[0041] As used herein, the terms "ligate," "ligation," and derivatives thereof generally refer to the process of covalently bonding two or more molecules to each other, e.g., covalently bonding two or more nucleic acid molecules to each other. In some embodiments, ligation includes joining nicks between adjacent nucleotides of a nucleic acid. In some embodiments, ligation includes forming a covalent bond between the end of a first nucleic acid molecule and the end of a second nucleic acid molecule. In some embodiments, ligation can include forming a covalent bond between the 5' phosphate group of one nucleic acid and the 3' hydroxyl group of a second nucleic acid, thereby forming a ligated nucleic acid molecule. Generally, for the purposes of the present disclosure, an amplified target sequence can be ligated to an adapter to generate an adapter-ligated amplified target sequence.
[0042] As used herein, "ligase" and derivatives thereof generally refer to any agent capable of catalyzing the ligation of two substrate molecules. In some embodiments, ligase includes an enzyme capable of catalyzing the joining of nicks between adjacent nucleotides of a nucleic acid. In some embodiments, ligase includes an enzyme capable of catalyzing the formation of a covalent bond between the 5' phosphate of one nucleic acid molecule and the 3' hydroxyl of another nucleic acid molecule, thereby forming a ligated nucleic acid molecule. Suitable ligases can include, but are not limited to, T4 DNA ligase, T4 RNA ligase, and E. coli DNA ligase.
[0043] As used herein, the term "ligation conditions" and derivatives thereof generally refer to conditions suitable for ligating two molecules to each other. In some embodiments, the ligation conditions are suitable for sealing nicks or gaps between nucleic acids. As used herein, the terms nick or gap are consistent with the use of the terms in the art. Typically, a nick or gap can be ligated in the presence of an enzyme such as ligase at an appropriate temperature and pH. In some embodiments, T4 DNA ligase can bind to a nick between nucleic acids at a temperature of about 70-72°C.
[0044] As used herein, the term "flow cell" refers to a chamber that includes a solid surface through which one or more fluid reagents can flow. Examples of flow cells and related fluid systems and detection platforms that can be readily used in the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), International Publication No. 04 / 018497, U.S. Patent No. 7,057,026, International Publication No. 91 / 06678, No. 07 / 123744, U.S. Patent Nos. 7,329,492, 7,211,414, 7,315,019, 7,405,281, and U.S. Patent Application Publication No. 2008 / 0108082.
[0045] As used herein, the term "amplicon", when used in reference to a nucleic acid, means a product that copies the nucleic acid, and this product has a nucleotide sequence that is the same as or complementary to at least a portion of the nucleotide sequence of the nucleic acid. An amplicon can be produced by any of a variety of amplification methods that use a nucleic acid or an amplicon thereof as a template, including, for example, polymerase extension, polymerase chain reaction (PCR), rolling circle amplification (RCA), ligation extension, or ligation chain reaction. An amplicon can be a nucleic acid molecule having a single copy of a particular nucleotide sequence (e.g., a PCR product) or multiple copies of a nucleotide sequence (e.g., a concatemer product of RCA). The first amplicon of a target nucleic acid is typically a complementary copy. Subsequent amplicons are copies made from the target nucleic acid or the first amplicon after the production of the first amplicon.
[0046] As used herein, the term "amplification site" refers to a site within or on an array at which one or more amplicons can be generated. The amplification site can be further configured to contain, hold, or attach at least one amplicon generated at that site.
[0047] As used herein, the term "array" refers to a collection of sites that can be distinguished from one another according to their relative positions. Different molecules at different sites of the array can be distinguished from one another according to the position of the sites within the array. An individual site of the array can contain one or more molecules of a particular type. For example, a site can contain a single target nucleic acid molecule having a particular sequence, or a site can contain several nucleic acid molecules having the same sequence (and / or its complementary sequence). The sites of the array can be different features located on the same substrate. Exemplary features include, but are not limited to, wells in the substrate, beads (or other particles) in or on the substrate, protrusions from the substrate, raised portions on the substrate, or channels within the substrate. The sites of the array can be separate substrates each having a different molecule. Different molecules attached to separate substrates can be identified according to the position of the substrates on the surface where the substrates associate, or according to the position of the substrates within a liquid or gel. Exemplary arrays in which separate substrates are disposed on a surface include, but are not limited to, those having beads in wells.
[0048] As used herein, the term "capacity", when used with respect to a site and nucleic acid material, means the maximum amount of nucleic acid material that can occupy the site. For example, the term can refer to the total number of nucleic acid molecules that can occupy the site under particular conditions. Other measurements can be used, including, for example, the total mass of nucleic acid material that can occupy the site under particular conditions or the total number of copies of a particular nucleotide sequence. Typically, the capacity of a site for a target nucleic acid is substantially equivalent to the capacity of the site for an amplicon of the target nucleic acid.
[0049] As used herein, the term "capture agent" refers to a material, chemical substance, molecule, or portion thereof that can attach to, hold, or bind to a target molecule (e.g., a target nucleic acid). Exemplary capture agents include a capture sequence (also referred to herein as a capture oligonucleotide) that is complementary to at least a portion of the target nucleic acid, a member of a receptor-ligand binding pair (e.g., avidin, streptavidin, biotin, lectin, carbohydrate, nucleic acid binding protein, epitope, antibody, etc.) that can bind to the target nucleic acid (or a linker moiety attached thereto), or a chemical reagent that can form a covalent bond with the target nucleic acid (or a linker moiety attached thereto), but is not limited thereto.
[0050] As used herein, the term "reporter moiety" can refer to any distinguishable tag, label, index, barcode, or group that enables determination of the composition, identity, and / or source of a target being investigated. In some embodiments, the reporter moiety can include an antibody that specifically binds to a protein. In some embodiments, the antibody may include a detectable label. In some embodiments, the reporter can include an antibody or affinity reagent labeled with a nucleic acid tag. In one embodiment, the nucleic acid is of a length sufficient to function as a substrate for a transpososome complex. In one embodiment, the nucleic acid tag can be detected, for example, via an epitope-based readout such as proximity ligation assay (PLA) or proximity extension assay (PEA), sequencing-based readout (Shahi et al. Scientific Reports volume7,Article number:44447,2017), or CITE-seq (Stoeckius et al. Nature Methods14:865-868,2017).
[0051] As used herein, the term "clone population" refers to a population of nucleic acids that are homogeneous with respect to a particular nucleotide sequence. The homogeneous sequence is typically at least 10 nucleotides in length, but can include longer lengths, for example, at least 50, 100, 250, 500, or 1000 nucleotides in length. A clone population can be derived from a single target nucleic acid or template nucleic acid. Typically, all of the nucleic acids in a clone population have the same nucleotide sequence. It will be understood that without departing from clonality, a small number of mutations (e.g., due to amplification artifacts) can occur.
[0052] As used herein, the term "unique molecular identifier" or "UMI" refers to a molecular tag that can be attached to a nucleic acid, which can be random, non-random, or semi-random. When incorporated into a nucleic acid, the unique molecular identifier (UMI) that is sequenced after amplification can be used to correct for subsequent amplification bias by directly counting the UMI.
[0053] As used herein, an "exogenous" compound, such as an exogenous enzyme, refers to a compound that is not normally or naturally found in a particular composition. For example, if a particular composition contains a cell lysate, an exogenous enzyme is an enzyme that is not normally or naturally found in the cell lysate.
[0054] As used herein, for example, "providing" in the context of a composition, article, nucleic acid, or nucleus means making the composition, article, nucleic acid, or nucleus, purchasing the composition, article, nucleic acid, or nucleus, or otherwise obtaining the compound, composition, article, or nucleus.
[0055] The term "and / or" means one or all of the listed elements, or any combination of two or more of the listed elements.
[0056] The terms "preferred" and "preferably" refer to embodiments of the present disclosure that may provide certain benefits under certain circumstances. However, in the same or other circumstances, other embodiments may be preferred. Further, the description of one or more preferred embodiments does not imply that other embodiments are not useful, nor is it intended to exclude other embodiments from the scope of the present disclosure.
[0057] The term "comprises" and its variations do not have a limiting meaning when these terms appear in the description and claims.
[0058] As used herein, when the terms "include", "includes", or "including" are used in this specification, embodiments similar to those described by the terms "consisting of" and / or "consisting essentially of" are also provided. It is understood that.
[0059] Unless otherwise stated, "a", "an", "the", and "at least one" are used interchangeably and mean one or more than one.
[0060] In this specification, the recitation of a numerical range by endpoints includes all numbers subsumed within that range (e.g., 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.80, 4, 5, etc.).
[0061] In any method disclosed herein that includes discrete steps, the steps may be performed in any executable order. Also, appropriately, any combination of two or more steps may be performed simultaneously.
[0062] References to "one embodiment", "an embodiment", "a particular embodiment", or "some embodiments" mean that a particular feature, structure, composition, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. Thus, appearances of such phrases in various places throughout this specification are not necessarily referring to the same embodiment of the present disclosure. Further, the particular features, structures, compositions, or characteristics may be combined in any suitable manner in one or more embodiments.
Brief Description of the Drawings
[0063] The following detailed description of exemplary embodiments of the present disclosure can be best understood when read in conjunction with the following drawings.
[0064]
Figure 1A
Figure 1B
[0065]
Figure 2
[0066]
Figure 3
[0067]
Figure 4
[0068]
Figure 5
[0069]
Figure 6
[0070]
Figure 7
[0071]
Figure 8
[0072]
Figure 9
[0073]
Figure 10
[0074]
Figure 11
[0075] The schematic diagrams are not necessarily to scale. Like numbers used in the drawings refer to like components, steps, etc. However, it will be understood that the use of numbers to refer to components of a given figure is not intended to limit the components in another figure labeled with the same number. Further, the use of different numbers to refer to components is not intended to indicate that components with different numbers cannot be the same or similar to components with other labels.
DETAILED DESCRIPTION OF THE INVENTION
[0076] The methods provided herein can be used to generate sequencing libraries from multiple single cells. In essence, single nucleus sequencing of transposon-accessible chromatin (sci-ATAC, U.S. Patent No. 10,059,989), single nucleus whole genome sequencing (U.S. Patent Application Publication No. 2018 / 0023119), single nucleus transcriptome sequencing (U.S. Provisional Patent Application No. 62 / 680,259 and Gunderson et al. (International Publication No. 2016 / 130704)), sci-HiC (Ramani et al., Nature Methods, 2017, 14:263-266), DRUG-seq (Ye et al., Nature Commun., 9, article number 4307), or any combination of analyses from DNA and proteins, e.g., sci-CAR (Cao et al., Science, 2018, 361(6409):1380-1385) and RNA and proteins, e.g., CITE-seq (Stoeckius et al., 2017, Nature Methods. 14(9):865-868), etc., but not limited thereto, any single nucleus or single cell library preparation method or sequencing method can be used. In one embodiment, the cell atlas experiment can be performed using reads limited to chromatin-accessible DNA, whole cell transcriptome, very informative, limited number of mRNAs, or combinations thereof.
[0077] Providing an isolated nucleus or cell
[0078] In one embodiment, the methods provided herein can include providing a nucleus isolated from a cell or cells (e.g., FIG. 1A, block 10, FIG. 3, block 30, FIG. 4, block 40, FIG. 6, block 600). The cell can be from any organism and can be from any cell type or any tissue of the organism. In one embodiment, the cell can be from a biopsy such as a tissue or a liquid biopsy. In one embodiment, the cell can be a germ cell, e.g., a cell obtained from an embryo. In one embodiment, the cell or nucleus can be from cancer or diseased tissue. In one embodiment, the cell or nucleus can be an immune cell such as a T cell or a B cell. In one embodiment, the cell can be various different cell types obtained from a single organism. In one embodiment, the various different cell types obtained from a single organism can include microbial cells such as prokaryotic cells and / or eukaryotic cells. In one embodiment, cells from different sources, e.g., different organisms and / or different tissues, are not combined at this stage. In one embodiment, cells from different sources, e.g., different organisms and / or different tissues, are combined at this stage.
[0079] In one embodiment, the plurality of cells can be a subset of a larger cell population. The subset can be separated from other cells based on differences in the size, shape, or presence or absence of identifiable molecules such as proteins or glycans on the surface of the cells. Methods for sorting cells are known in the art and include fluorescence-activated cell sorting, magnetic-activated cell sorting, and microfluidic cell sorting.
[0080] The method can further include dissociating the cells and / or isolating the nucleus. In one embodiment, conditions are used that maintain the chromatin present in the nucleus. In one embodiment, the nucleosomes present in the nucleus are depleted. Methods for depleting nucleosomes are known to those of skill in the art (U.S. Patent Application Publication No. 2018 / 002311).
[0081] Many different single-cell library preparation methods are known in the art. (Hwang et al., Experimental & Molecular Medicine, vol. 50, Article number: 96 (2018)), including, but not limited to, the Drop-seq method, the Seq-well method, and the single-cell combinatorial indexing (“sci-”) method. Companies that provide single-cell products and related technologies include, but are not limited to, 10X Genomics, Takara biosciences, BD biosciences, Biorad, 1cellbio, IsoPlexis, CellSee, NanoCellect, and Dolomite Bio. SCI-seq is a methodological framework that uniquely labels the nucleic acid content of a large number of single cells or single nuclei using split-pool barcoding. Typically, the number of nuclei or cells can be at least two. The upper limit depends on the actual limitations of the equipment used in other steps of the methods described herein (e.g., multi-well plates, number of indexes). The number of nuclei or cells that can be used is not intended to be limiting and can reach billions. For example, in one embodiment, the number of nuclei or cells can be 1,000,000,000 or less, 100,000,000 or less, 10,000,000 or less, 1,000,000 or less, 100,000 or less, 10,000 or less, 1,000 or less, 500 or less, or 50 or less. In one embodiment, the number of nuclei or cells can be at least 50, at least 500, at least 1,000, at least 10,000, at least 100,000, at least 1,000,000, at least 10,000,000, at least 100,000,000, or at least 1,000,000,000.
[0082] In these embodiments using isolated nuclei, the nuclei can be obtained by extraction and fixation. Optionally, and preferably, the method for obtaining isolated nuclei does not include enzymatic treatment.
[0083] In one embodiment, nuclei are isolated from individual cells that are adhesives or suspensions. Methods for isolating nuclei from individual cells are known to those skilled in the art. Nuclei are typically isolated from cells present in a tissue. Methods for obtaining isolated nuclei typically include preparing the tissue, isolating the nuclei from the prepared tissue, and then fixing the nuclei. In one embodiment, all steps are performed on ice.
[0084] In one embodiment, tissue preparation includes flash-freezing the tissue in liquid nitrogen and then reducing the size of the tissue to pieces having a diameter of 1 mm or less. The tissue can be sized down by subjecting it to either a mincing force or a blunt force. Mincing can be accomplished with a blade for cutting the tissue into small pieces. Applying a blunt force can be accomplished by crushing the tissue with a hammer or similar object, and the composition resulting from the crushed tissue is called a powder.
[0085] Nuclear isolation can be achieved by incubating the pellet or powder in cell lysis buffer for at least 1 minute to 20 minutes, such as 5 minutes, 10 minutes, or 15 minutes. Useful buffers are those that promote cell lysis but maintain nuclear integrity. Examples of cell lysis buffers include 10 mM Tris-HCl, pH 7.4, 10 mM NaCl, 3 mM MgCl2, 0.1% IGEPAL CA-630, 1% SUPERase In RNase inhibitor (20 U / μL, Ambion), and 1% BSA (20 mg / mL, NEB). Standard nuclear isolation methods often use one or more exogenous compounds, such as exogenous enzymes, to assist in the isolation. Examples of useful enzymes that may be present in the cell lysis buffer include protease inhibitors, lysozyme, proteinase K, surfactants, lysostaphin, zymolyase, cellulase, protease, or glycanase, etc. (Islam et al. Micromachines (Basel), 2017, 8(3):83; www.sigmaaldrich.com / life-science / biochemicals / biochemical-products.html?TablePage=14573107), but are not limited thereto. In one embodiment, one or more exogenous enzymes are not present in the cell lysis buffer useful in the methods described herein. For example, the exogenous enzyme is (i) not added to the cells prior to mixing the cells with the lysis buffer, (ii) not present in the cell lysis buffer prior to mixing with the cells, (iii) not added to the mixture of the cells and the cell lysis buffer, or combinations thereof. One of ordinary skill in the art will recognize that the concentrations of these components can be varied to some extent without reducing the usefulness of the cell lysis buffer for isolating nuclei. The extracted nuclei are then purified by one or more rounds of washing with nuclear buffer. Examples of nuclear buffers include , 10 mM Tris-HCl, pH 7.4, 10 mM NaCl, 3 mM MgCl2, 1% SUPERase In RNase inhibitor (20 U / μL, Ambion), and 1% BSA (20 mg / mL, NEB). Similar to the cell lysis buffer, exogenous enzymes may also not be present in the nuclear buffer used in the methods of the present disclosure. One of ordinary skill in the art will recognize that the concentrations of these components can be varied to some extent without reducing the usefulness of the nuclear buffer for isolating nuclei. One of ordinary skill in the art will recognize that BSA and / or a surfactant may be useful in the buffers used for nuclear isolation.
[0086] The isolated nuclei can be fixed by exposure to a cross-linking agent. Useful examples of cross-linking agents include, but are not limited to, paraformaldehyde and formaldehyde. Paraformaldehyde can be at a concentration of 1% - 8%, such as 4%. Formaldehyde can be at a concentration of 30% - 45%, such as 37%. Treatment of the nuclei with a cross-linking agent can include adding the cross-linking agent to the nuclear suspension and incubating at 0°C. Other methods of fixation include, but are not limited to, methanol fixation. Optionally and preferably, after fixation, washing in the nuclear buffer is performed.
[0087] The isolated fixed nuclei can be immediately aliquoted and snap-frozen in liquid nitrogen for later use. When prepared for use after freezing, the thawed nuclei can be permeabilized, for example, with 0.2% Triton X-100 on ice for 3 minutes and sonicated briefly to reduce nuclear aggregation.
[0088] Conventional tissue nucleus extraction techniques typically incubate tissues with tissue-specific enzymes (e.g., trypsin) at a high temperature (e.g., 37°C) for 30 minutes to several hours, and then lyse the cells with cell lysis buffer. The nucleus isolation method described herein has several advantages. That is, (1) no artificial enzymes are introduced and all steps are performed on ice. This reduces potential perturbations to the cellular state (e.g., chromatin organization state, or transcriptome state). (2) This new method has been validated across most tissue types, including diseased samples such as brain, lung, kidney, spleen, heart, cerebellum, and tumor tissue. Compared to conventional tissue nucleus extraction techniques that use different enzymes for different tissue types, the new technique can potentially reduce bias when comparing cellular states from different tissues. (3) This new method also reduces costs and increases efficiency by removing the enzyme treatment step. (4) Compared to other nucleus extraction techniques (e.g., Dounce tissue grinder), this new technique is more robust for different tissue types (e.g., the Dounce method requires optimizing the Dounce cycle for different tissues) and can process large sample pieces with high throughput (e.g., the Dounce method is limited by the size of the grinder).
[0089] Optionally, the isolated nuclei may not contain nucleosomes or may be subjected to conditions that deplete the nucleosomes in the nuclei to produce nucleosome-depleted nuclei.
[0090] Insertion of Universal Sequences
[0091] The methods provided herein include inserting one or more universal sequences into nucleic acids present in nuclei or cells. In one embodiment, the incorporation of one or more universal sequences occurs prior to the distribution of subsets (Figure 1A, block 11, Figure 1B, block 110), and in other embodiments, the incorporation of one or more universal sequences occurs after the distribution of subsets (Figure 3, block 32, Figure 4, blocks 42, 45). In some embodiments, an index may also be combined with the universal sequence or may be associated with the cell or nucleus as an optional step separate from the insertion of one or more universal sequences. Optional indexing of the nucleus or cell can occur before or after the insertion of the universal sequence (Figure 1A, block 12). In one embodiment, an index is added to the sample prior to the distribution of a subset of nuclei or cells (Figure 1A, block 13). In some embodiments, an index is added to a plurality of samples prior to the distribution of a subset of nuclei or cells (Figure 1A, block 13).
[0092] In one embodiment, a transposome complex is used. The transposome complex binds to a transposase recognition site and can insert the transposase recognition site into a target nucleic acid in the nucleus in a process sometimes called "tagging." In some such insertion events, a single strand of the transposase recognition site can be transferred to the target nucleic acid. Such a strand is referred to as the "transfer strand." In one embodiment, the transposome complex includes a dimer transposase having two subunits and two non - contiguous transposon sequences. In another embodiment, the transposase includes a dimer transposase having two subunits and a non - contiguous transposon sequence. In one embodiment, the 5' end of one or both strands of the transposase recognition site can be phosphorylated.
[0093] Some embodiments can include the use of hyperactive Tn5 transposase and Tn5 transposase recognition sites (Goryshin and Reznikoff, J. Biol. Chem., 273:7367 (1998)), or MuA transposase and Mu transposase recognition sites (Mizuuchi, K., Cell, 35:785, 1983, Savilahti, H et al., EMBO J., 14:4893, 1995) that include R1 and R2 end sequences. The Tn5 mosaic end (ME) sequences can also be used by those skilled in the art.
[0094] As further examples of transposition systems that can be used with certain embodiments of the compositions and methods provided herein, there are Staphylococcus aureus Tn552 (Colegio et al., J. Bacteriol., 183:2384-8, 2001, Kirby C et al., Mol. Microbiol., 43:173-86, 2002), Ty1 (Devine and Boeke, Nucleic Acids Res., 22:3765-72, 1994, and WO 95 / 23875), transposon Tn7 (Craig, NL, Science. 271:1512, 1996, Craig, NL, Curr Top Reviews in Microbiol Immunol., 204:27-48, 1996), Tn10 and IS10 (Kleckner N, et al., Curr Top Microbiol Immunol., 204:49-82, 1996), Mariner transposase (Lampe D J, et al., EMBO J, 15:5470-9, 1996), Tc1 (Plasterk R H, Curr.Topics Microbiol.Immunol., 204:125-43, 1996), P element (Gloor, G B, Methods Mol.BiBiol, 260:97-114, 2004), Tn3 (Ichikawa and Ohtsubo, J Biol.Chem.265:18829-32, 1990), bacterial insertion sequences (Ohtsubo and Sekine, Curr.Top.Microbiol.Immunol.204:1-26, 1996), retroviruses (Brown et al., Proc Natl Acad Sci USA, 86:2525-9, 1989), and yeast retrotransposons (Boeke and Corces, Annu Rev Microbiol.43:403-34, 1989). Other examples include IS5, Tn10, Tn903, IS911, and modified forms of transposase family enzymes (Zhang et al., (2009) PLoS Genet.5:e1000689.Epub Oct 16, 2009, Wilson C. et al. (2007) J.Microbiol.Methods 71:332-5).
[0095] Other examples of integrases that can be used with the methods and compositions provided herein include retroviral integrases and integrase recognition sequences of such retroviral integrases, for example, integrases from HIV-1, HIV-2, SIV, PFV-1, RSV.
[0096] Transposon arrays useful in the methods and compositions described herein are described in U.S. Patent Application Publication No. 2012 / 0208705, U.S. Patent Application Publication No. 2012 / 0208724, and International Publication No. 2012 / 061832. In some embodiments, the transposon array comprises a first transposase recognition site and a second transposase recognition site.
[0097] Some transposome complexes useful herein comprise a transposase having two transposon arrays. In some such embodiments, the two transposon arrays are not linked to each other; in other words, the transposon arrays are not contiguous with each other. Examples of such transposomes are known in the art (see, for example, U.S. Patent Application Publication No. 2010 / 0120098).
[0098] In one embodiment, tagging is used to generate a target nucleic acid that includes different universal sequences at each end (e.g., a universal primer binding site such as A14 at one end and a universal primer binding site such as B15 at the other end). This can be achieved by using two types of transposome complexes, each of which includes a different nucleotide sequence that is part of the transfer strand. The universal sequences can serve multiple purposes. For example, without intending to be limiting, the universal sequences can serve as complementary sequences for hybridization in subsequent amplification steps to add another nucleotide sequence (e.g., an index), can serve as a site for annealing of a universal primer (e.g., a sequencing primer for read 1 or read 2) for sequencing, or can serve as a "landing pad" in subsequent steps for annealing a nucleotide sequence that can be used as a primer to add another nucleotide sequence, such as an index, to the target nucleic acid.
[0099] In some embodiments, the transpososome complex comprises a transposon array nucleic acid that joins two transposase subunits to form a "loop complex" or "loop transpososome". In one example, the transpososome comprises a dimeric transposase and a transposon array. The loop complex can ensure that the transposon is inserted into the target DNA while maintaining the order information of the original target DNA without fragmenting the target DNA. As will be appreciated, the loop structure may insert a desired nucleic acid sequence, such as a universal sequence, into the target nucleic acid while maintaining the physical connectivity of the target nucleic acid. In some embodiments, the transposon array of the loop transpososome complex can include a fragmentation site so that a transpososome complex can be created that fragments the transposon array to include two transposon arrays. Such a transpososome complex is useful for ensuring that neighboring target DNA fragments into which the transposon is inserted receive a combination of barcodes that can be unambiguously assembled at a later stage of the assay. In one embodiment, a combination of indexes is added after insertion of one or more universal sequences into the target nucleic acid.
[0100] In one embodiment, fragmentation of the nucleic acid is achieved by using a fragmentation site present in the nucleic acid. Typically, the fragmentation site is introduced into the target nucleic acid by using a transpososome complex. In one embodiment, after fragmentation of the nucleic acid fragments, the transposase remains bound to the nucleic acid fragments such that nucleic acid fragments derived from the same genomic DNA molecule remain physically linked (Adey et al., 2014, Genome Res., 24: 2041 - 2049, Amini S. et al. (2014) Nat Genet 46: 1343 - 1349). For example, a looped transpososome complex may contain fragmentation sites. Fragmentation sites can be used to cleave physical associations, but cannot be used to cleave the informational associations between index sequences incorporated into the target nucleic acid. Cleavage may be effected by biochemical, chemical, or other means. In some embodiments, the fragmentation site may contain nucleotides or nucleotide sequences that can be fragmented by various means. Examples of fragmentation sites include restriction endonuclease sites, at least one ribonucleotide cleavable by RNAse, nucleotide analogs cleavable in the presence of specific chemical agents, diol bonds cleavable by periodate treatment, disulfide groups cleavable by chemical reducing agents, cleavable moieties that can be subjected to photochemical cleavage, and peptides cleavable by peptidase enzymes or other suitable means, but are not limited thereto (see, for example, U.S. Patent Application Publication No. 2012 / 0208705, U.S. Patent Application Publication No. 2012 / 0208724, and International Publication No. 2012 / 061832). In one embodiment, the transposase remains bound to the nucleic acid fragment and maintains the physical association between nucleic acid fragments derived from the same genomic DNA molecule until removal by the use of appropriate conditions such as the addition of a protein denaturant (e.g., SDS) or a chelating agent (e.g., EDTA). This type of approach allows for the derivation of contiguous information by capturing continuously ligated and transposed target nucleic acids (U.S. Patent Application Publication No. 2019 / 0040382). Contiguous information can be preserved by maintaining the association of adjacent template nucleic acid fragments within the target nucleic acid using a transposase.
[0101] Instead of translocation, target nucleic acids can be obtained by fragmentation. Fragmentation of the primary nucleic acid from the sample can be achieved in any order by enzymatic, chemical, or mechanical methods, and then adapters are added to the ends of the fragments. Examples of enzymatic fragmentation include CRISPR and Talen-like enzymes, and enzymes that unwind DNA (e.g., helicases) that can create single-stranded regions where DNA fragments can hybridize and initiate extension or amplification. For example, helicase-based amplification can be used (Vincent et al., 2004, EMBO Rep., 5(8):795-800). In one embodiment, the extension or amplification is initiated using random primers. Examples of mechanical fragmentation include nebulization or sonication.
[0102] Fragmentation of the primary nucleic acid by mechanical means results in fragments having a heterogeneous mixture of blunt ends, 3' overhang ends, and 5' overhang ends. Thus, it is desirable to repair the fragment ends using methods known in the art, for example, to generate ends that are optimal for adding adapters to blunt sites. In certain embodiments, the fragment ends of the nucleic acid population are blunt ends. More specifically, the fragment ends are blunt ends and are phosphorylated. The phosphate moiety can be introduced by enzymatic treatment, for example, using polynucleotide kinase.
[0103] In one embodiment, the fragmented nucleic acid is prepared using overhang nucleotides. For example, a single overhang nucleotide can be added by a specific type of activity, such as Taq polymerase or Klenow exo-minus polymerase, which has template-independent terminal transferase activity and adds a single deoxynucleotide, such as adding the nucleotide "A" to the 3' end of a DNA molecule. Using such an enzyme, a single nucleotide "A" can be added to the 3' end of the blunt ends of each strand of a double-stranded nucleic acid fragment. Thus, by reaction with Taq or Klenow exo-minus polymerase, "A" can be added to the 3' end of each strand of the end-repaired double-stranded target fragment, while the adapter can be a T construct having a compatible "T" overhang present at the 3' end of each region of the double-stranded nucleic acid of the universal adapter. In one example, terminal deoxynucleotidyl transferase (TdT) can be used to add multiple "T" nucleotides (Swift Biosciences, Ann Arbor, MI). This type of end modification also prevents self-ligation of both the vector and the target so that there is a bias to form target nucleic acids having the same adapter at each end.
[0104] The primary nucleic acid can be DNA, RNA, or a DNA / RNA hybrid. In embodiments where the primary nucleic acid is RNA, incorporating one or more universal sequences into the nucleic acid present in the nucleus or cell typically involves converting the RNA to DNA. Various methods can be used, but in some embodiments, conventional methods used to produce cDNA are included. For example, a primer having a polyT sequence at the 3' end and an adapter upstream of the polyT sequence can be annealed to the mRNA molecule and extended using reverse transcriptase. This results in a one-step conversion of mRNA to DNA and, optionally, a one-step conversion of a universal sequence to the 3' end. In one embodiment, the primer may also include one or more index sequences. In one embodiment, random primers are used.
[0105] Non-coding RNAs can also be converted to DNA and optionally modified to contain universal sequences using various methods. For example, an adapter can be added using a first primer that includes a random sequence and a template-switch primer, and either primer can include a universal sequence adapter. A reverse transcriptase having terminal transferase activity can be used to effect the addition of non-template nucleotides to the 3' end of the synthetic strand, and the template-switch primer includes nucleotides that anneal to the non-template nucleotides added by the reverse transcriptase. An example of a useful reverse transcriptase is Moloney-mouse leukemia virus reverse transcriptase. In certain embodiments, for use in template-switching, the SMARTer™ reagent (catalog number 634926) available from Takara Bio USA, Inc. is used to add a universal sequence to non-coding RNA and, optionally, to mRNA. Optionally, the template-switch primer can be used with RNA in combination with a primer having a poly-T sequence to add universal sequences to both ends of the DNA target nucleic acid produced from the RNA.
[0106] Allocation of subsets
[0107] The methods provided herein include distributing isolated nuclei or a subset of cells into a plurality of compartments (Figure 1A, block 13, Figure 1B, block 115, Figure 3, block 31, Figure 4, blocks 41, 44). The methods can include multiple partitioning steps of dividing a population of isolated nuclei or cells (also referred to herein as a pool) into subsets. Typically, a subset of isolated nuclei or cells, e.g., a subset present in a plurality of compartments, is indexed with a compartment-specific index and then pooled. Thus, the methods typically include at least one "split and pool" step of obtaining pooled isolated nuclei or isolated cells, distributing them, and adding a compartment-specific index, and the number of "split and pool" steps can depend on the number of different indexes added to the target nucleic acid. Each initial subset of nuclei or cells prior to indexing can be unique and different from other subsets. For example, each first subset can be from a unique sample such as a unique organism or a unique tissue. After indexing, the subsets can be pooled, divided into subsets, and pooled again as needed until a sufficient number of indexes are added to the target nucleic acid. This process results in a combinatorial indexing as described herein by assigning an index or combination of indexes unique to each single cell or single nucleus. After completion of indexing, e.g., after addition of one, two, three, or more indexes, the isolated nuclei or cells can be lysed. In some embodiments, addition of the index and lysis can occur simultaneously.
[0108] The subset, and thus the number of nuclei or cells present within each compartment, can be at least 1. In one embodiment, the number of nuclei or cells present within the subset is 100,000,000 or less, 10,000,000 or less, 1,000,000 or less, 100,000 or less, 10,000 or less, 4,000 or less, 3,000 or less, 2,000 or less, 1,000 or less, 500 or less, or 50 or less. In one embodiment, the number of nuclei or cells present within the subset can be from 1 to 1,000, from 1,000 to 10,000, from 10,000 to 100,000, from 100,000 to 1,000,000, from 1,000,000 to 10,000,000, or from 10,000,000 to 100,000,000. In one embodiment, the number of nuclei or cells present within each subset is approximately equal. The number of nuclei or cells present within the subset, and thus the number of nuclei or cells within each compartment, is based in part on the desire to reduce index collisions, where a collision is the presence of two nuclei or cells having the same combination of indices that end up in the same compartment in this step of the method. Methods for distributing nuclei or cells into subsets are known to those of skill in the art and routine. Fluorescence-activated cell sorting (FACS) cytometry can be used, although in some embodiments, the use of simple dilution is preferred. In one embodiment, FACS cytometry is not used. Optionally, staining, such as DAPI (4’,6-diamidino-2-phenylindole) staining, can be used to gate and enrich nuclei of different ploidies. The staining can also be used to identify single cells from doublets during sorting.
[0109] The number of partitions in the partitioning process (and subsequent index addition) can depend on the format used. For example, the number of partitions can be 2 to 96 partitions (when using a 96-well plate), 2 to 384 partitions (when using a 384-well plate), or 2 to 1536 partitions (when using a 1536-well plate). In one embodiment, multiple plates can be used. Examples of partitions include, but are not limited to, wells, droplets, and microfluidic partitions. In one embodiment, each partition can be a droplet. When the type of partition used is a droplet containing two or more nuclei or cells, any number of droplets can be used, such as at least 10,000, at least 100,000, at least 1,000,000, or at least 10,000,000 droplets. A subset of isolated nuclei or cells is typically indexed within the partition prior to pooling.
[0110] Combinatorial indexing
[0111] The methods provided herein include adding a partition-specific index to nuclei or cells present in a sample (Figure 1B, block 112), or adding a partition-specific index to a subset of isolated nuclei or cells that are distributed into different partitions (e.g., Figure 1A, block 14, Figure 3, block 32, Figure 4, blocks 42 and 45, Figure 6, block 601). In some embodiments, the universal sequence can also be incorporated with the index. The index sequence, also referred to as a tag or barcode, is useful as a marker characteristic of the partition in which a particular nucleic acid is present. Thus, in some embodiments, the index is a nucleic acid sequence tag that is bound to each of the target nucleic acids present in a particular partition, the presence of which is used to indicate or identify the partition in which a population of nuclei or cells is present at a particular stage of the method.
[0112] In one embodiment, multiple indexes are added. The incorporation of each index occurs with a single split and pool indexing. Split and pool barcoding one, two, three, or more times results in a targeted nucleic acid with single, double, triple, or multiple (e.g., quadruple or higher) indexes.
[0113] The index can be added to one or both ends of the target nucleic acid. For example, a modified target nucleic acid having two or more indexes can include different indexes at each end (an example is shown in FIG. 5A). In FIG. 5A, the target nucleic acid 55 is modified to include four distinct indexes, two indexes (51 and 52) at one end, and two indexes (53 and 54) at the other end. In other embodiments, the modified target nucleic acid can include indexes grouped at one or both ends (an example is shown in FIG. 5B). In FIG. 5B, the target nucleic acid 56 is modified to include four distinct indexes (51, 52, 53, and 54) at each end. A set of indexes present at one end of the target nucleic acid can be referred to as a "consecutive index." In one embodiment, the consecutive index has no nucleotides between each index. In other embodiments, one or more nucleotides can be present between one or more of the indexes of the consecutive index. As described herein, consecutive indexes can be useful in identifying members of a library having a specific index set. For example, consecutive indexes can facilitate the enrichment of library members derived from the same cell.
[0114] The index sequence can be of any suitable number, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides in length. A four-nucleotide tag provides the possibility of multiplexing 256 samples on the same array, and a six-base tag enables the processing of 4096 samples on the same array.
[0115] In one embodiment, the index is added, for example by a transpososome complex, after the universal array has been integrated into the nuclear or cellular DNA nucleic acid. The integration of the index sequence can essentially use any combination of ligation, extension, hybridization, adsorption, specific or non-specific interaction of primers, or amplification to include one, two, or more steps. In one embodiment, the index is added during cDNA synthesis. In one embodiment, the index is added through tagging. The nucleotide sequence added to one or both ends of the target nucleic acid may also include other useful sequences such as one or more universal arrays and / or unique molecular identifiers.
[0116] Various methods can be used to add an index to a nucleic acid containing a universal array, and it is not intended to limit the method of adding the index. In one embodiment, the target nucleic acid has different universal arrays (e.g., A14 at one end and B15 at the other end) at each end, and those skilled in the art will recognize that specific sequences can be added to one or both ends of the target nucleic acid. The universal array added by the transpososome complex can be used as a "landing pad" in a subsequent step of annealing a nucleotide sequence that can be used as a primer to add another nucleotide sequence, such as another index and / or another universal array, to the target nucleic acid. For example, in one embodiment, the integration of the index sequence includes ligating primers to one or both ends of the nucleic acid. The ligation of the primers can be assisted by the presence of the universal array at each end of the target nucleic acid. An example of a primer is double hairpin ligation. Double ligation can be ligated to one end, or preferably both ends, of the target nucleic acid.
[0117] In one embodiment, blunt end ligation can be used. In another embodiment, the target nucleic acid is prepared using a single overhang nucleotide by the activity of a specific type of DNA polymerase, such as Taq polymerase, or a template-independent terminal transferase activity that adds one or more deoxynucleotides, such as deoxyadenosine (A), to the 3' end of the target nucleic acid, such as Klenow exo-minus polymerase. In some cases, the overhang nucleotide is two or more bases. Using such an enzyme, a single nucleotide "A" can be added to the 3' end, which is the blunt end of each strand of the target nucleic acid. Thus, by reaction with Taq or Klenow exo-minus polymerase, "A" can be added to the 3' end of each strand of the double-stranded target fragment, while the additional sequence added to each end of the target nucleic acid can include a complementary "T" overhang present at the 3' end of each region of the double-stranded nucleic acid to be added. This terminal modification also prevents self-ligation of the nucleic acid such that there is a bias to form an indexed target nucleic acid adjacent to the sequence added in this embodiment.
[0118] In one embodiment, incorporation of the index is performed by an exponential amplification reaction such as PCR. The universal sequence present at the ends of the target nucleic acid serves as a primer and can be used for annealing of sequences that can be extended in the amplification reaction.
[0119] The index and other useful sequences can be added in a single step or in multiple steps. For example, the index and any other useful sequences can be added by ligation or extension, or a two-step method can be used that includes, for example, ligating a universal sequence and then amplifying to further modify the universal sequence to include the index and any other useful sequences.
[0120] In one embodiment, the addition of sequences during the indexing step adds universal sequences useful for immobilization and / or sequencing of the target nucleic acid. In another embodiment, the indexed target nucleic acid can be further processed to add universal sequences useful for immobilization and sequencing of the target nucleic acid. Those skilled in the art will recognize that in embodiments where the compartments are droplets, the sequences for immobilizing the nucleic acid fragments are optional. In one embodiment, the incorporation of universal sequences useful for fragment immobilization and sequencing involves ligating the same universal adapter (also referred to as a "mismatch adapter", the general features of which are described in U.S. Patent Nos. 7,741,463 (Gormley et al.) and 8,053,192 (Bignell et al.)) to the 5' and 3' ends of the indexed nucleic acid fragment. In one embodiment, the universal adapter contains all the sequences necessary for sequencing, including sequences for immobilizing the indexed nucleic acid fragment on the array.
[0121] The resulting indexed fragments collectively provide a library of nucleic acids that can be immobilized and then sequenced. The term library, also referred to herein as a sequencing library, refers to an aggregate of nucleic acid fragments from a single nucleus or single cell that contains various combinations of known universal sequences and indexes at the 3' and 5' ends. The library can include, for example, accessible DNA, whole genomes, or whole transcriptomes, nucleic acids that encode a particular protein, or nucleic acids from combinations thereof, and can be used for performing sequencing.
[0122] Indexed nucleic acid fragments can be subjected to a selection condition for a predetermined size range such as 150 to 400 nucleotides, such as 150 to 300 nucleotides. The obtained indexed nucleic acid fragments are pooled and optionally subjected to a cleanup process to improve the purity of the DNA molecules by removing at least a portion of the unincorporated universal adapter or primer. Any suitable cleanup process such as electrophoresis, size exclusion chromatography, etc. may be used. In some embodiments, solid phase reversible immobilization paramagnetic beads may be used to separate the desired DNA molecules from the unbound universal adapter or primer and select nucleic acids based on size. Solid phase reversible immobilization paramagnetic beads are commercially available from Beckman Coulter (Agencourt AMPure XP), Thermo Fisher (MagJet), Omega Biotech (Mag-Bind), Promega Beads, and Kapa Biosystems (Kapa Pure Beads).
[0123] A non-limiting, exemplary embodiment of the present disclosure is shown in FIG. 1A. In this embodiment, the method includes providing a plurality of nuclei or cells (FIG. 1A, block 10). The plurality of nuclei or cells can be from a sample or a plurality of samples. The method further includes incorporating one or more universal sequences into the nucleic acids present in the nuclei or cells (FIG. 1A, block 11). Optionally, the method can also include associating an index with the nuclei or cells (see, e.g., nuclear or cell hashing, International Publication No. WO 2020 / 180778), and in one embodiment, by associating, an index can be added to the nucleic acids (FIG. 1A, block 12). In one embodiment, two different universal sequences are added such that ultimately a target nucleic acid having different universal sequences at each end is obtained. The method further includes distributing a subset of the nuclei or cells, incorporating universal sequences into the nucleic acids located therein, and optionally incorporating at least one index into a plurality of compartments (FIG. 1, block 13). Indexing the nucleic acids present in each compartment (FIG. 1A, block 14), and then pooling the nuclei or cells (FIG. 1A, block 15). After the addition of a single index, the library of nucleic acids within the nuclei or cells can be further processed to prepare for sequencing (FIG. 1A, block 16). However, in some preferred embodiments, it is desirable to add a second, third, or more indices. In one embodiment, the addition of each index can include a "split and pool" step in which indexing occurs after splitting, e.g., distributing a subset of the nuclei or cells into a plurality of compartments (FIG. 1A, block 13), indexing the nucleic acids present within each compartment (FIG. 1A, block 14), and then pooling the nuclei or cells (FIG. 1A, block 15). The "split and pool" step can result in adding an index to only one or both ends of the nucleic acids present in the nuclei or cells. After the addition of the last index, the library of nucleic acids within the nuclei or cells can be pooled and further processed to prepare for sequencing (which can be whole-genome sequencing or targeted sequencing) (FIG. 1A, block 16).
[0124] Another non-limiting exemplary embodiment of the present disclosure is shown in FIG. 1B. In this embodiment, the method includes first providing a plurality of samples to be processed in parallel (FIG. 1B, block 110). The method includes incorporating one or more universal sequences into the nucleic acids present in the nuclei or cells (FIG. 1B, block 111), followed by adding an index to the nucleic acids (FIG. 1B, block 112), wherein the index added to each sample is unique and can be used as a sample index to identify the nucleic acids derived from a particular sample. In one embodiment, two different universal sequences are added such that ultimately a target nucleic acid having different universal sequences at each end is obtained. The method further includes pooling the nuclei or cells (FIG. 1B, block 113). In one embodiment, after the addition of one index, the library of nucleic acids within the nuclei or cells can be further processed to prepare for sequencing (FIG. 1B, block 114). However, in some preferred embodiments, it is desirable to add a second, third, or more indices. In one embodiment, the addition of each index can include a "split and pool" step in which indexing occurs after splitting, for example, by distributing a subset of the nuclei or cells into a plurality of compartments (FIG. 1B, block 115), indexing the nucleic acids present within each compartment (FIG. 1B, block 116), and then pooling the nuclei or cells (FIG. 1B, block 117). The "split and pool" step can result in adding an index to only one or both ends of the nucleic acids present in the nuclei or cells. After the addition of the last index, the library of nucleic acids within the nuclei or cells can be pooled and further processed to prepare for sequencing (which can be whole-genome sequencing or targeted sequencing) (FIG. 1B, block 118).
[0125] Another non-limiting, exemplary embodiment of the present disclosure is shown in FIG. 2. In this embodiment, the method involves using tagging to incorporate two universal sequences into the nucleic acids present in the nuclei or cells and performing three subsequent indexing steps (FIG. 2A). One transpososome complex 21 includes a universal sequence 23 (e.g., A14), and another transpososome complex 22 includes a universal sequence 24 (B15). Insertion of the universal sequences into the nucleic acids occurs for a bulk of nuclei or cells. FIG. 2A also shows the result of the insertion of the two universal sequences 23 and 24 into the target nucleic acid 25. The plurality of nuclei or cells are distributed into different compartments, and a polynucleotide 26 containing an index is added to one side of the nucleic acid 25 by ligation using nucleotides complementary to one of the universal sequences (e.g., A14) (FIG. 2B). The plurality of nuclei or cells are pooled and then distributed into different compartments, and a different polynucleotide 27 containing a second index is added to the other side of the nucleic acid 25 by ligation using nucleotides complementary to the other universal sequence (e.g., B15) (FIG. 2C). The plurality of nuclei or cells containing the double-indexed nucleic acids are pooled and then distributed into different compartments, and then subjected to a PCR amplification reaction in which a polynucleotide 28 containing a third index is added to one side of the nucleic acid 25 and a polynucleotide 29 containing a fourth index is added to one side of the nucleic acid 25 (FIG. 2D). After the addition of the last index, the library of nucleic acids within the nuclei or cells can be pooled and further processed to prepare for sequencing (which can be whole-genome sequencing or targeted sequencing).
[0126] Yet another non-limiting exemplary embodiment of the present disclosure is shown in FIG. 3. In this embodiment, the method includes providing a plurality of nuclei or cells (FIG. 3, block 30). The method further includes distributing a subset of the nuclei or cells into a plurality of compartments (FIG. 3, block 31). The nucleic acids present in the nuclei or cells of each compartment are modified by incorporation of an index and / or a universal sequence (FIG. 3, block 32). In another embodiment, the nucleic acids present in the nuclei or cells of each compartment are modified by incorporation of the same universal sequence (e.g., tagging using a transposon having the same universal sequence), followed by addition of a compartment-specific index. The nuclei or cells are then pooled (FIG. 3, block 33). After addition of the index and / or universal sequence, the library of nucleic acids within the nuclei or cells can be further processed to prepare for sequencing (FIG. 3, block 34). However, in some preferred embodiments, it is desirable to add a second, third, or more indices. Optionally, a universal sequence can also be added. The addition of each index can include a "split and pool" step in which indexing occurs after splitting, e.g., distributing a subset of the nuclei or cells into a plurality of compartments (FIG. 3, block 31), indexing the nucleic acids present within each compartment (FIG. 3, block 32), and then pooling the nuclei or cells (FIG. 3, block 33). The "split and pool" step can result in addition of an index to only one or both ends of the nucleic acids present in the nuclei or cells. After addition of the last index, the library of nucleic acids within the nuclei or cells is pooled and further processed to prepare for sequencing (which can be comprehensive sequencing or targeted sequencing) (FIG. 3, block 34).
[0127] A further non-limiting exemplary embodiment of the present disclosure is shown in FIG. 4. In this embodiment, the method includes the analysis of RNA. A plurality of nuclei or cells are provided (FIG. 4, block 40), which can be obtained from a sample or a plurality of samples. A subset of the nuclei or cells is distributed into a plurality of compartments (FIG. 4, block 41). Optionally, the method may also include associating an index with the nuclei or cells (e.g., see nuclear or cell hashing, WO 2020 / 180778) or with the nucleic acids prior to distribution. The nucleic acids present in the nuclei or cells of each compartment are modified using reverse transcriptase and an index and / or a universal sequence is inserted (FIG. 4, block 42), and then the nuclei or cells are pooled (FIG. 4, block 43). The method further includes distributing a subset of the nuclei or cells into a plurality of compartments (FIG. 4, block 44). The nucleic acids present in the nuclei or cells of each compartment are modified by the insertion of a different index and / or universal sequence (FIG. 4, block 45), and then the nuclei or cells are pooled (FIG. 4, block 46). After the addition of the index and / or universal sequence, the library of nucleic acids within the nuclei or cells can be further processed to prepare for sequencing (FIG. 4, block 47). However, in some preferred embodiments, it is desirable to add a third, fourth, or more indices. Optionally, a universal sequence can also be added. The addition of each index can include a "split and pool" step in which indexing occurs after splitting, e.g., distributing a subset of the nuclei or cells into a plurality of compartments (FIG. 4, block 44), indexing the nucleic acids present within each compartment (FIG. 4, block 45), and then pooling the nuclei or cells (FIG. 4, block 46). The "split and pool" step can result in adding an index to only one or both ends of the nucleic acids present in the nuclei or cells. After the addition of the last index, the library of nucleic acids within the nuclei or cells is pooled and further processed to prepare for sequencing (which can be whole sequencing or targeted sequencing) (FIG. 4, block 47).
[0128] Preparation of a Fixed Sample for Sequencing
[0129] Methods for attaching indexed fragments from one or more sources to a substrate are known in the art. In one embodiment, the indexed fragments are enriched using a plurality of capture arrays having specificity for the indexed fragments, and the capture arrays can be immobilized on the surface of a solid substrate. For example, the capture array can include a first member of a binding pair (e.g., P5'), and a second member (P5) of the binding pair is immobilized on the surface of the solid substrate. Similarly, methods for amplifying the immobilized indexed fragments include, but are not limited to, bridge amplification and binding equilibrium exclusion. Methods for immobilization and amplification prior to sequencing are described, for example, in Bignell et al. (U.S. Patent No. 8,053,192), Gunderson et al. (International Publication No. 2016 / 130704), Shen et al. (U.S. Patent No. 8,895,249), and Pipenburg et al. (U.S. Patent No. 9,309,502).
[0130] Pooled samples can be immobilized during preparation for sequencing. Sequencing can be performed as a single molecule array or amplified prior to sequencing. Amplification can be performed using one or more immobilized primers. The immobilized primers can be, for example, on a plane or in a lawn on a pool of beads. The pool of beads can be isolated in an emulsion having a single bead in each "compartment" of the emulsion. At a concentration of only one template per "compartment", only a single template is amplified on each bead.
[0131] As used herein, the term "solid-phase amplification" refers to any nucleic acid amplification reaction that is performed on or in association with a solid support such that all or a portion of the amplification product is immobilized on the solid support upon formation. Specifically, this term encompasses solid-phase polymerase chain reaction (solid-phase PCR) and solid-phase isothermal amplification, which are reactions similar to standard solution-phase amplification except that one or both of the forward and reverse amplification primers are immobilized on the solid support. Solid-phase PCR is directed to systems such as emulsions where one primer is immobilized on beads and the other is in the free solution, or colony formation in a solid-phase gel matrix where one primer is immobilized on the surface and the other is in the free solution.
[0132] In some embodiments, the solid support comprises a patterned surface. A "patterned surface" refers to the arrangement of different regions within or on the exposed layer of the solid support. For example, one or more regions can be features where one or more amplification primers are present. This feature can be separated by interstitial regions where no amplification primer is present. In some embodiments, the pattern can be in an x-y format of features in rows and columns. In some embodiments, the pattern can be a repeating sequence of features and / or interstitial regions. In some embodiments, the pattern can be a random sequence of features and / or interstitial regions. Exemplary patterned surfaces that can be used in the methods and compositions described herein are described in U.S. Patent Nos. 8,778,848, 8,778,849, and 9,079,148, and U.S. Patent Application Publication No. 2014 / 0243224.
[0133] In some embodiments, the solid support comprises an array of wells or depressions on the surface. This can be manufactured in a manner generally known in the art using a variety of techniques including, but not limited to, photolithography, stamping techniques, molding techniques, and microetching techniques. As understood in the art, the techniques used depend on the composition and shape of the array substrate.
[0134] Features within the patterned surface can be wells of an array of wells (e.g., microwells or nanowells) on another suitable solid support comprising a patterned covalent gel such as glass, silicon, plastic, or poly(N-(5-azidoacetamylpentyl)acrylamide-co-acrylamide) (PAZAM, see, e.g., U.S. Patent Application Publication No. 2013 / 184796, International Publication No. 2016 / 066586, and No. 2015 / 002813). This process creates gel pads for use in sequencing, which can be stable over multiple cycles of sequencing operation. Covalently binding a polymer to the wells is useful for maintaining the gel in the structured features over the lifetime of the structured substrate among various applications. However, in many embodiments, the gel need not be covalently bound to the wells. For example, under some conditions, a silane-free acrylamide (SFA, see, e.g., U.S. Patent No. 8,563,477) that is not covalently bound to any part of the structured substrate can be used as the gel material.
[0135] In certain other embodiments, the structured substrate is made by patterning a solid support material using wells (e.g., microwells or nanowells), coating the patterned support with a gel material (e.g., PAZAM, SFA, or a chemically modified variant thereof), such as an azidated form of SFA (azido-SFA), and polishing the gel-coated support, e.g., by chemical or mechanical polishing, thereby retaining the gel within the wells but removing or inactivating substantially all of the gel from the interstitial regions of the surface of the structured substrate between the wells. A primer nucleic acid can be attached to the gel material. Next, a solution of indexed fragments can be contacted with the polished substrate such that individual indexed fragments can be seeded into individual wells via interaction with the primer attached to the gel material, but since the gel material is absent or inactive, the target nucleic acid does not occupy the interstitial regions. Amplification of the indexed fragments will be limited to the wells because the absence or inactivation of the gel in the interstitial regions prevents outward movement of the growing nucleic acid colonies. The process is conveniently manufacturable, scalable, and utilizes conventional micro- or nano-fabrication methods.
[0136] The present disclosure encompasses “solid-phase” amplification methods in which only one amplification primer is immobilized (the other primers are typically present in free solution), but in one embodiment it is desirable for the solid support to be provided with both immobilized forward and reverse primers. In practice, since the amplification process requires an excess of primers to maintain amplification, there will be “multiple” identical forward primers and / or “multiple” identical reverse primers immobilized on the solid support. References herein to forward and reverse primers should be construed to include “multiple” such primers unless the context indicates otherwise.
[0137] As will be appreciated by those skilled in the art, any given amplification reaction requires at least one type of forward primer and at least one type of reverse primer that are specific to the template being amplified. However, in certain embodiments, the forward and reverse primers may include template-specific portions of the same sequence and may have exactly the same nucleotide sequence and structure (including any non-nucleotide modifications). In other words, solid-phase amplification can be performed using only one type of primer, and such single-primer methods are encompassed within the scope of the present disclosure. Other embodiments may use forward and reverse primers that include the same template-specific sequence but differ in some other structural feature. For example, one type of primer may include non-nucleotide modifications that are not present in the other.
[0138] In all embodiments of the present disclosure, the primer for solid-phase amplification is preferably immobilized by a single-point covalent bond to the solid support near or at the 5' end of the primer, allowing the template-specific portion of the primer to freely anneal to its cognate template and the 3'-hydroxyl group that does not include primer extension. Any suitable covalent bonding means known in the art can be used for this purpose. The selected attachment chemistry depends on the nature of the solid support and any derivatization or functionalization applied thereto. The primer itself may include a portion that may be a non-nucleotide chemical modification to facilitate attachment. In certain embodiments, the primer may include a sulfur-containing nucleophile such as phosphorothioate or thiophosphate at the 5' end. In the case of a solid-supported polyacrylamide hydrogel, this nucleophile binds to the bromoacetamide groups present in the hydrogel. A more specific means of binding the primer and template to the solid support is via a 5'-phosphorothioate bond to a hydrogel consisting of polymerized acrylamide and N-(5-bromoacetamido-pentyl)acrylamide (BRAPA), as described in International Publication No. 05 / 065814.
[0139] Certain embodiments of the present disclosure utilize a solid support comprising an inert substrate or matrix (e.g., a glass slide, polymer beads, etc.) that has been "functionalized" by the application of a layer or coating of an intermediate material that includes reactive groups enabling covalent attachment to biomolecules such as polynucleotides. Examples of such supports include, but are not limited to, polyacrylamide hydrogels supported on an inert substrate such as glass. In such embodiments, a biomolecule (e.g., a polynucleotide) may covalently attach directly to the intermediate material (e.g., a hydrogel), which intermediate material may itself non-covalently attach to the substrate or matrix (e.g., a glass substrate). The term "covalent attachment to a solid support" should be construed as appropriate to encompass this type of arrangement.
[0140] Pooled samples may be amplified on beads, each bead containing forward and reverse amplification primers. In certain embodiments, a library of indexed fragments is used to prepare a clustered array of nucleic acid colonies by solid-phase amplification, more specifically solid-phase isothermal amplification, similar to that described in U.S. Patent Application Publication No. 2005 / 0100900, U.S. Patent No. 7,115,400, International Publication Nos. 00 / 18957 and 98 / 44151. The terms "cluster" and "colony" are used interchangeably herein and refer to discrete sites on a solid support that include multiple identical immobilized nucleic acid strands and multiple identical immobilized complementary nucleic acid strands. The term "clustered array" refers to an array formed from such clusters or colonies. In this context, the term "array" should not be understood to require an ordered arrangement of clusters.
[0141] The term "solid phase" or "surface" is used to mean either a flat surface to which a primer is attached, such as a glass, silica or plastic microscope slide, or a similar flow cell device, or beads, to which one or two primers are attached and on which the beads are amplified, or an array of beads on a surface after the beads have been amplified.
[0142] The clustered arrays can be adjusted using a process of thermal cycling as described in WO 98 / 44151, or a process in which the temperature is maintained constant and cycles of extension and denaturation are performed using changes in reagents. Such isothermal amplification methods are described in WO 02 / 46456 and US 2008 / 0009420 A1. Due to the lower temperatures useful in isothermal processes, this is particularly preferred in some embodiments.
[0143] It will be understood that any of the amplification methods described herein or generally known in the art can be used with universal or target-specific primers to amplify immobilized DNA fragments. Suitable methods for amplification include, but are not limited to, polymerase chain reaction (PCR), strand displacement amplification (SDA), transcription-mediated amplification (TMA), and nucleic acid sequence-based amplification (NASBA) as described in US Pat. No. 8,003,354. Using the above amplification methods, one or more nucleic acids of interest can be amplified. For example, immobilized DNA fragments can be amplified using PCR such as multiplex PCR, SDA, TMA, NASBA. In some embodiments, primers specifically directed to the polynucleotide of interest are included in the amplification reaction.
[0144] Other methods suitable for amplifying polynucleotides may include oligonucleotide extension and ligation, rolling circle amplification (RCA) (Lizardi et al., Nat. Genet. 19:225-232 (1998)), and oligonucleotide ligation assay (OLA) (see generally U.S. Patent Nos. 7,582,420, 5,185,243, 5,679,524, and 5,573,907; European Patent Nos. 0 320 308 (B1), 0 336 731 (B1), 0 439 182 (B1); International Publication Nos. 90 / 01069, 89 / 12696, and 89 / 09835). It will be understood that these amplification methods can be designed to amplify immobilized DNA fragments. For example, in some embodiments, the amplification method may include a ligation probe amplification or oligonucleotide ligation assay (OLA) reaction containing primers specifically directed to the nucleic acid of interest. In some embodiments, the amplification method may include a primer extension ligation reaction containing primers specifically directed to the nucleic acid of interest. Non-limiting examples of primer extension and ligation primers that can be specifically designed to amplify the nucleic acid of interest include primers used in the GoldenGate assay, as exemplified by U.S. Patent Nos. 7,582,420 and 7,611,869 (Illumina, San Diego, California).
[0145] DNA nanoblocks can also be used in combination with the methods and compositions described herein. Methods for making and using DNA nanoblocks for genome sequencing can be found, for example, in U.S. Patent Nos. 7,910,354; 2009 / 0264299; 2009 / 0011943; 2009 / 0005252; 2009 / 0155781; 2009 / 0118488, and in, for example, Drmanac et al., 2010, Science 327(5961):78 - 81. Briefly, after fragmentation of genomic library DNA, adapters are ligated to the fragments, the adapter-ligated fragments are circularized by ligation with circle ligase, and rolling circle amplification is performed (described in Lizardi et al., 1998, Nat. Genet. 19:225 - 232 and U.S. Patent Application Publication No. 2007 / 0099208(A1)). The extended concatemer structure of the amplicon promotes coiling, thereby generating compact DNA nanoballs. The DNA nanoballs can be captured on a substrate to form an ordered or patterned array, preferably such that the distance between each nanoball is maintained, thereby enabling sequencing of individual DNA nanoballs. In some embodiments, adapter ligation, amplification, and digestion are performed prior to circularization to create a head-to-tail construct having several genomic DNA fragments separated by adapter sequences.
[0146] Exemplary isothermal amplification methods that can be used in the methods of the present disclosure include, for example, multiple displacement amplification (MDA) exemplified by isothermal strand displacement nucleic acid amplification as exemplified by Dean et al., Proc. Natl. Acad. Sci. USA 99:5261-66 (2002), or, for example, U.S. Patent No. 6,214,587, but are not limited thereto. Other non-PCR-based methods that can be used in the present disclosure include, for example, strand displacement amplification (SDA) described by Walker et al., Molecular Methods for Virus Detection, Academic Press, 1995, U.S. Patent Nos. 5,455,166 and 5,130,238, and Walker et al., Nucl. Acids Res. 20:1691-96 (1992), or, for example, hyperbranched strand displacement amplification described by Lage et al., Genome Res. 13:294-307 (2003). Isothermal amplification methods can be used, for example, with strand displacement Phi 29 polymerase or large fragments of Bst DNA polymerase for 5'->3' exo for random primer amplification of genomic DNA. The use of these polymerases takes advantage of their high processivity and strand displacement activity. Due to high processivity, the polymerase can produce fragments 10-20 kb in length. As described above, polymerases with low processivity and polymerases with strand displacement activity such as Klenow polymerase can be used to produce smaller fragments under isothermal conditions. Further descriptions of the amplification reactions, conditions, and components are detailed in the disclosure of U.S. Patent No. 7,670,810.
[0147] Another polynucleotide amplification method useful in the present disclosure is tagged PCR, which uses a population of two-domain primers having a random 3' region following a 5' region, as described, for example, in Grothues et al., Nucleic Acids Res. 21(5):1321-2 (1993). The first round of amplification is performed to allow multiple starts on heat-denatured DNA based on individual hybridization from randomly synthesized 3' regions. Due to the nature of the 3' region, the starting sites are thought to be random throughout the genome. Thereafter, unbound primers may be removed and further replication may be performed using primers complementary to a constant 5' region.
[0148] In some embodiments, isothermal amplification can be performed using binding equilibrium exclusion amplification (KEA), also known as exclusion amplification (ExAmp). The nucleic acid libraries of the present disclosure can be made using a method that includes reacting amplification reagents to produce a plurality of amplification sites from individual target nucleic acids seeded at the sites, each containing a substantially clonal population of amplicons. In some embodiments, the amplification reaction proceeds until a sufficient number of amplicons are produced to fill the volume of each amplification site. Thus, when a previously seeded site is filled to capacity, it inhibits the target nucleic acid from landing and amplifying at that site, thereby producing a clonal population of amplicons at that site. In some embodiments, apparent clonality can be achieved even if the amplification site is not filled to capacity before a second target nucleic acid reaches that site. Under some conditions, amplification of the first target nucleic acid can proceed to the point where a sufficient number of copies are made to effectively exceed or overwhelm the production of copies from the second target nucleic acid transported to that site. For example, in embodiments using a bridge amplification process on circular features less than 500 nm in diameter, after 14 cycles of exponential amplification for the first target nucleic acid, contamination from the second target nucleic acid at the same site was determined to generate an insufficient number of contaminating amplicons to adversely affect sequence synthesis analysis on an Illumina sequencing platform.
[0149] In some embodiments, the amplification sites in the array can be, but need not necessarily be, completely clonal. Rather, in some applications, the individual amplification sites can be predominantly occupied by amplicons from the first indexed fragment and can also have low levels of contaminating amplicons from a second target nucleic acid. The array can have one or more amplification sites with low levels of contaminating amplicons as long as the contamination level does not have an unacceptable impact on subsequent use of the array. For example, when the array is used for detection applications, an acceptable level of contamination is a level that does not affect the signal-to-noise ratio or resolution of the detection technique in an unacceptable way. Thus, apparent clonality generally relates to a particular use or application of the array produced by the methods described herein. Exemplary levels of contamination acceptable at an individual amplification site for a particular application include, but are not limited to, up to 0.1%, 0.5%, 1%, 5%, 10%, or 25% contaminating amplicons. The array can include one or more amplification sites having these exemplary levels of contaminating amplicons. For example, up to 5%, 10%, 25%, 50%, 75%, or even 100% of the amplification sites within the array can contain contaminated amplicons. It will be appreciated that in an array or other collection of sites, at least 50%, 75%, 80%, 85%, 90%, 95%, or 99% or more of the sites can be clonal or can appear clonal.
[0150] In some embodiments, binding equilibrium exclusion can occur when the process occurs at a sufficiently fast rate to effectively exclude the occurrence of another event or process. Taking as an example the creation of a nucleic acid array in which sites of the array are randomly seeded with indexed fragments from a solution and copies of the indexed fragments are produced in an amplification process to fill each of the seeding sites to capacity. According to the binding equilibrium exclusion method of the present disclosure, the seeding and amplification processes can proceed simultaneously under conditions where the amplification rate exceeds the seeding rate. Thus, the relatively fast rate at which copies are made at sites seeded by the first target nucleic acid effectively excludes the second nucleic acid from seeding that site for amplification. The binding equilibrium exclusion amplification method can be practiced as described in detail in the disclosure of U.S. Patent Application Publication No. 2013 / 0338042.
[0151] Binding equilibrium exclusion can utilize a relatively slow rate for initiating amplification (e.g., a slow rate for making the first copy of an indexed fragment) versus a relatively fast rate for making subsequent copies of the indexed fragment (or the first copy of the indexed fragment). In the example of the previous paragraph, binding equilibrium exclusion results from a relatively slow rate of seeding of indexed fragments (e.g., relatively slow diffusion or transport) versus the relatively fast rate at which amplification occurs to fill sites with copies of the indexed fragment seed. In another exemplary embodiment, binding equilibrium exclusion can result from a delay in the formation of the first copy of an indexed fragment that seeds a site (e.g., a delay or slow activation) versus the relatively fast rate at which subsequent copies are made to fill the site. In this example, several different indexed fragments may be seeded at an individual site (e.g., several indexed fragments may be present at each site prior to amplification). However, since the formation of the first copy of any given indexed fragment can be randomly activated, the average rate of first copy formation is relatively slow compared to the rate at which subsequent copies are generated. In this case, several different indexed fragments may be seeded at an individual site, but binding equilibrium exclusion allows only one of those indexed fragments to be amplified. More specifically, when the first indexed fragment is activated for amplification, the site is rapidly filled to capacity with its copies, thereby preventing copies of the second indexed fragment from being made at the site.
[0152] In one embodiment, the method is performed simultaneously to (i) transport indexed fragments to the amplification site at an average transport rate and (ii) amplify the indexed fragments at the amplification site at an average amplification rate that exceeds the average transport rate (U.S. Patent No. 9,169,513). Thus, in such an embodiment, exclusion of binding equilibrium can be achieved by using a relatively slow transport rate. For example, a lower concentration results in a slower transport rate, so a sufficiently low concentration of indexed fragments can be selected to achieve the desired average transport rate. Alternatively or additionally, the presence of a high viscosity solution and / or a molecular crowding reagent in the solution can be used to decrease the transport rate. Examples of useful molecular crowding reagents include, but are not limited to, polyethylene glycol (PEG), ficoll, dextran, or polyvinyl alcohol. Exemplary molecular crowding reagents and formulations are described in U.S. Patent No. 7,399,590, which is incorporated herein by reference. Another factor that can be adjusted to achieve the desired transport rate is the average size of the target nucleic acid.
[0153] The amplification reagent can include additional components that promote amplicon formation and, in some cases, increase the rate of amplicon formation. One example is recombinase. Recombinase can promote amplicon formation by enabling iterative invasion / elongation. More specifically, recombinase can facilitate the invasion of index fragments by polymerase and the elongation of primers by polymerase using the indexed fragments as templates for amplicon formation. This process can be repeated as a chain reaction where the amplicons produced from each round of invasion / elongation function as templates in subsequent rounds. Since denaturation cycles (e.g., by heating or chemical denaturation) are not required, this process can be performed more rapidly than standard PCR. Thus, recombinase-mediated amplification can be performed isothermally. To promote amplification, it is desirable to include ATP, or other nucleotides (or, in some cases, their non-hydrolyzable analogs) in the recombinase-mediated amplification reagent. A mixture of recombinase and single-strand binding (SSB) protein is particularly useful because SSB can further enhance amplification. Representative formulations for recombinase-mediated amplification include those commercially available as the TwistAmp kit from TwistDx (Cambridge, UK). Useful components and reaction conditions for recombinase-mediated amplification reagents are described in U.S. Patent Nos. 5,223,414 and 7,399,590.
[0154] Another example of a component that can be included in an amplification reagent to promote amplicon formation and, in some cases, increase the rate of amplicon formation is a helicase. Helicases can promote amplicon formation by enabling the chain reaction of amplicon formation. Since a denaturation cycle (e.g., by heating or chemical denaturation) is not required, this process can be performed more rapidly than standard PCR. Thus, helicase-promoted amplification can be performed isothermally. A mixture of helicase and single-strand binding (SSB) protein is particularly useful because SSB can further enhance the amplification. A representative formulation for helicase-promoted amplification is commercially available as the IsoAmp kit from Biohelix (Beverly, Massachusetts). Further, examples of useful formulations containing helicase protein are described in U.S. Patent Nos. 7,399,590 and 7,829,284.
[0155] Yet another example of a component that can be included in an amplification reagent to promote amplicon formation and, in some cases, increase the rate of amplicon formation is a origin-binding protein.
[0156] Sequencing method
[0157] Following the attachment of the indexed fragments to the surface, the array of immobilized and amplified indexed fragments is determined. Sequencing can be either whole sequencing or targeted sequencing. Whole sequencing can be used when the entire sequence of each cell or nucleus present in the library is desired. Examples of applications using whole sequencing include, but are not limited to, whole genome sequencing, whole transcriptome sequencing, and ATAC sequencing. Targeted sequencing can be used when information regarding biological characteristics is desired. In one embodiment, targeted sequencing can be used to identify a subpopulation of cells or nuclei, or a subset of the genome, transcriptome, proteome, or any combination thereof, as described in detail herein.
[0158] Sequencing can be performed using any suitable sequencing technology, and methods for determining the sequence of immobilized and amplified indexed fragments, such as strand resynthesis, are known in the art and are described, for example, in Bignell et al. (U.S. Patent No. 8,053,192), Gunderson et al. (International Publication No. 2016 / 130704), Shen et al. (U.S. Patent No. 8,895,249), and Pipenburg et al. (U.S. Patent No. 9,309,502).
[0159] The methods described herein can be used in conjunction with various nucleic acid sequencing methods. Particularly applicable techniques are those in which the nucleic acids are attached at fixed positions within an array such that their relative positions do not change and the array is imaged repeatedly. For example, embodiments in which images are obtained in different color channels corresponding to different labels used to distinguish one nucleotide base type from another are particularly applicable. In some embodiments, the process of determining the nucleotide sequence of the indexed fragments can be an automated process. Preferred embodiments include sequencing by synthesis ("SBS") technology.
[0160] SBS technology generally involves the enzymatic extension of a nascent nucleic acid strand by iterative addition of nucleotides to a template strand. In conventional methods of SBS, a single nucleotide monomer can be provided to the target nucleotide in the presence of polymerase in each delivery. However, in the methods described herein, multiple types of nucleotide monomers can be provided to the target nucleic acid in the presence of polymerase during delivery.
[0161] In one embodiment, the nucleotide monomer includes a locked nucleic acid (LNA) or a bridged nucleic acid (BNA). The use of LNA or BNA in the nucleotide monomer increases the hybridization strength between the nucleotide monomer and the sequencing primer sequence present on the immobilized indexed fragment.
[0162] SBS can use nucleotide monomers having a terminator portion or nucleotide monomers lacking a terminator portion. Examples of methods using nucleotide monomers without a terminator include, for example, pyrosequencing and sequencing using γ-phosphate-labeled nucleotides as described in more detail herein. In methods using nucleotide monomers without a terminator, the number of nucleotides added in each cycle is generally variable and depends on the template sequence and the mode of nucleotide delivery. In SBS technology utilizing nucleotide monomers having a terminator portion, the terminator can be effectively irreversible under the sequencing conditions used as in conventional Sanger sequencing utilizing dideoxynucleotides, or the terminator can be reversible as in the sequencing method developed by Solexa (now Illumina, Inc.).
[0163] The SBS technique can use nucleotide monomers with a label part or nucleotide monomers lacking a label part. Therefore, incorporation events can be detected based on properties of the label such as fluorescence of the label, properties of the nucleotide monomers such as molecular weight or charge, by-products of nucleotide incorporation such as the release of pyrophosphate, etc. In embodiments where two or more different nucleotides are present in the sequencing reagent, the different nucleotides may be distinguishable from each other, or two or more different labels may be distinguishable under the detection technique used. For example, different nucleotides present in the sequencing reagent can have different labels, and they can be distinguished using an appropriate optical system exemplified by the sequencing method developed by Solexa (now Illumina).
[0164] Preferred embodiments include pyrosequencing technology. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) when a specific nucleotide is incorporated into a nascent strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M. and Nyren, P. (1996) "Real-time DNA sequencing using detection of pyrophosphate release." Analytical Biochemistry 242(1), 84-9, Ronaghi, M. (2001) "Pyrosequencing sheds light on DNA sequencing." Genome Res., 11(1), 3-11, Ronaghi, M., Uhlen, M. and Nyren, P. (1998) "A sequencing method based on real-time pyrophosphate." Science 281(5375), 363, U.S. Patent Nos. 6,210,891, 6,258,568 and 6,274,320). In pyrosequencing, the released PPi can be detected by its immediate conversion to adenosine triphosphate (ATP) by ATP sulfurylase, and the level of the generated ATP is detected via the photons generated by luciferase. The nucleic acid to be sequenced can be attached to features in an array, and the array can be imaged to capture the chemiluminescent signals produced by incorporating nucleotides into the features of the array. An image can be obtained after treating the array with a specific nucleotide type (e.g., T, C, or G). The images obtained after the addition of each nucleotide type differ with respect to which features within the array are detected. These differences within the image reflect the different sequence contents of the features on the array. However, the relative positions of each feature remain unchanged within the image. The images can be stored, processed, and analyzed using the methods described herein.For example, the images obtained after processing the array with each different nucleotide type can be processed in the same manner as exemplified herein for images obtained from different detection channels for reversible terminator-based sequencing methods.
[0165] In another exemplary type of SBS, cycle sequencing is achieved by stepwise addition of reversible terminator nucleotides that include cleavable or photo-bleachable dye labels, as described, for example, in WO 04 / 018497 and U.S. Patent No. 7,057,026. This approach has been commercialized by Solexa (now Illumina) and is also described in WO 91 / 06678 and WO 07 / 123,744. The availability of fluorescently labeled terminators whose fluorescence labels can be reversed allows for efficient cyclic reversible termination (CRT) sequencing. The polymerase can also co-operate to efficiently incorporate and extend from these modified nucleotides.
[0166] In some reversible terminator-based sequencing embodiments, the label does not substantially inhibit extension under SBS reaction conditions. However, the detection label may be removable, for example, by cleavage or degradation. Images can be captured after incorporation of the label into the arrayed nucleic acid features. In certain embodiments, each cycle involves the simultaneous delivery of four different nucleotide types to the array, and each nucleotide type has spectrally distinct labels. Next, four images can be obtained, each using a detection channel selective for one of the four different labels. Alternatively, the different nucleotide types can be added sequentially, and an image of the array can be obtained during each addition step. In such embodiments, each image shows nucleic acid features incorporating a particular type of nucleotide. Because the sequence content of each feature is different, different features are present in, or absent from, the various images. However, the relative positions of the features remain unchanged within the image. Images obtained from such reversible terminator-SBS methods can be stored, processed, and analyzed as described herein. Following the image capture step, the label can be removed, and the reversible terminator moiety can be removed for subsequent cycles of nucleotide addition and detection. Removing the label after detection in a particular cycle and prior to subsequent cycles has the advantage of reducing background signal and crosstalk between cycles. Examples of useful labels and removal methods are described herein.
[0167] In certain embodiments, some or all of the nucleotide monomers may include reversible terminators. In such embodiments, the reversible terminator / cleavable fluorophore may include a fluorophore attached to the ribose moiety via a 3'-ester linkage (Metzker, Genome Res. 15:1767-1776 (2005)). Other approaches have separated the chemical moiety of the terminator from the fluorescent label (Ruparel et al., Proc Natl Acad Sci USA 102:5932-7 (2005)). Ruparel et al. describe the development of a reversible terminator that uses a small amount of 3'-allyl group to block elongation but can be easily unblocked by treatment with a palladium catalyst for a short time. The fluorophore is attached to the group via a photocleavable linker that can be easily cleaved by exposure to long-wavelength UV light for 30 seconds. Thus, either disulfide reduction or photocleavage can be used as the cleavable linker. Another approach to reversible termination is the use of a natural termination following placement of a bulky dye on the dNTP. The presence of the charged bulky dye on the dNTP can act as an effective terminator via steric and / or electrostatic hindrance. The presence of one incorporation event prevents further binding unless the dye is removed. Cleavage of the dye removes the fluorophore and effectively reverses the termination. Examples of modified nucleotides are also described in U.S. Patent Nos. 7,427,673 and 7,057,026.
[0168] Additional exemplary SBS systems and methods that can be used with the methods and systems described herein are described in U.S. Patent Application Publication Nos. 2007 / 0166705, 2006 / 0188901, 2006 / 0240439, 2006 / 0281109, 2012 / 0270305, and 2013 / 0260372, U.S. Patent No. 7,057,026, and International Publication Nos. 05 / 065814, U.S. Patent Application Publication No. 2005 / 0100900, and International Publication Nos. 06 / 064199 and 07 / 010,251.
[0169] Some embodiments can use the detection of four different nucleotides using fewer than four different labels. For example, SBS can be implemented using the methods and systems described in U.S. Patent Publication No. 2013 / 0079232, which is incorporated herein by reference. As a first example, nucleotide type pairs can be detected at the same wavelength, but based on the difference in intensity for one member of the pair, or based on a change to one member of the pair (e.g., through chemical modification, photochemical modification, or physical modification) that causes a distinct signal to appear or disappear as compared to the signal detected for the other member of the pair. As a second example, three of the four different nucleotide types can be detected under certain conditions, while the fourth nucleotide type has no detectable label or is minimally detected under those conditions (e.g., minimal detection by background fluorescence). Incorporation of the first three nucleotide types into a nucleic acid can be determined based on the presence of their corresponding signals, and incorporation of the fourth nucleotide type into a nucleic acid can be determined based on the absence or minimal detection of any signal. As a third example, one nucleotide type can include a label that is detected in two different channels, while other nucleotide types are detected in one or fewer channels. The three exemplary configurations described above are not considered mutually exclusive and can be used in various combinations.An exemplary embodiment combining all three examples is a first nucleotide type detected in a first channel (e.g., dATP having a label detected in the first channel when excited by a first excitation wavelength), a second nucleotide type detected in a second channel (e.g., dCTP having a label detected in the second channel when excited by a second excitation wavelength), a third nucleotide type detected in both the first and second channels (e.g., dTTP having at least one label detected in both channels when excited by the first and / or second excitation wavelengths), and a fourth nucleotide type without a label (e.g., dGTP without a label) that is not detected or minimally detected in any channel, using a fluorescence-based SBS method.
[0170] Furthermore, as described in incorporated U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single channel. In such a so-called one-dye sequencing method, the first nucleotide type is labeled, but the label is removed after the first image is generated, and the second nucleotide type is labeled only after the first image is generated. The third nucleotide type retains its label in both the first and second images, and the fourth nucleotide type remains unlabeled in both images.
[0171] Some embodiments can use ligation-based sequencing techniques. Such techniques use DNA ligase to incorporate oligonucleotides and identify the incorporation of such oligonucleotides. Oligonucleotides typically have different labels that correlate with the identity of specific nucleotides in the sequence to which the oligonucleotide hybridizes. As with other SBS methods, after treating an array of nucleic acid sequences with labeled sequencing reagents, an image can be obtained. Each image shows a nucleic acid feature incorporating a particular type of label. Because the sequence content of each feature is different, different features may or may not be present in different images, but the relative positions of the features remain unchanged within the image. Images obtained from ligation-based sequencing methods can be stored, processed, and analyzed as described herein. Exemplary SBS systems and methods that can be used with the methods and systems described herein are described in U.S. Patent Nos. 6,969,488, 6,172,218, and 6,306,597.
[0172] Some embodiments can use nanopore sequencing (Deamer, D.W. & Akeson, M., "Nanopores and nucleic acids: prospects for ultrarapid sequencing.", Trends Biotechnol., 18, 147-151 (2000), Deamer, D. and D. Branton, "Characterization of nucleic acids by nanopore analysis", Acc. Chem. Res. 35:817-825 (2002), Li, J., M. Gershow, D. Stein, E. Brandin, and J.A. Golovchenko, "DNA molecules and configurations in a solid-state nanopore microscope" Nat. Mater. 2:611-615 (2003)). In such embodiments, the indexed fragment passes through the nanopore. The nanopore can be a synthetic pore such as α-hemolysin or a biomembrane protein. As the indexed fragment passes through the nanopore, each base pair can be identified by measuring fluctuations in the electrical conductance of the pore. (U.S. Patent No. 7,001,792, Soni, G.V. & Meller, "A. Progress toward ultrafast DNA sequencing using solid-state nanopores.", Clin. Chem. 53, 1996-2001 (2007), Healy, K., "Nanopore-based single-molecule DNA analysis.", Nanomed., 2, 459-481 (2007), Cockroft, S.L., Chu, J., Amorin, M. & Ghadiri, M.R., "A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution.", J. Am Chem. Soc. 130, 818-820 (2008).Data obtained from nanopore sequencing can be stored, processed, and analyzed as described herein. Specifically, the data can be processed as an image according to the exemplary processing of the optical and other images described herein.
[0173] Some embodiments can use methods that include real-time monitoring of DNA polymerase activity. Incorporation of nucleotides can be detected, for example, via a fluorescence resonance energy transfer (FRET) interaction between a fluorophore-containing polymerase and a γ-phosphate-labeled nucleotide as described in U.S. Patent Nos. 7,329,492 and 7,211,414, or incorporation of nucleotides can be detected using, for example, zero-mode waveguides as described in U.S. Patent No. 7,315,019, and fluorescence nucleotide analogs and engineered polymerases as described in, for example, U.S. Patent Nos. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082. Illumination can be restricted to a zeptoliter-scale volume around the surface-tethered polymerase such that incorporation of fluorescently labeled nucleotides can be observed with low background (Levene, M. J. et al., “Zero-mode waveguides for single-molecule analysis at high concentrations.” Science, 299, 682-686 (2003), Lundquist, P. M. et al., “Parallel confocal detection of single molecules in real time.” Opt. Lett. 33, 1026-1028 (2008), Levene, M. J. et al., “Zero-mode waveguides for single-molecule analysis at high concentrations.” Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008)). Images obtained from such methods can be stored, processed, and analyzed as described herein.
[0174] Some SBS embodiments include the detection of protons released upon incorporation of nucleotides into an extension product. For example, sequencing based on the detection of released protons can use an electrical detector and related technologies commercially available from Ion Torrent (Guilford, CT, a subsidiary of Life Technologies), or the sequencing methods and systems described in U.S. Patent Application Publication Nos. 2009 / 0026082, 2009 / 0127589, 2010 / 0137143, and 2010 / 0282617. The methods described herein for amplifying a target nucleic acid using binding equilibrium exclusion can be readily applied to the substrates used for detecting protons. More specifically, the methods described herein can be used to generate a clonal population of amplicons used for detecting protons.
[0175] The SBS methods described above can be advantageously implemented in multiplex format such that multiple different indexed fragments are manipulated simultaneously. In certain embodiments, the differently indexed fragments can be processed in a common reaction vessel or on the surface of a particular substrate. This enables facile delivery of sequencing reagents, removal of unreacted reagents, and multiplex detection of incorporation events. In embodiments using surface-bound target nucleic acids, the indexed fragments can be in array format. In array format, the indexed fragments can typically be bound to the surface in a spatially distinguishable manner. The indexed fragments can be bound by direct covalent attachment, attachment to beads or other particles, or binding to a polymerase or other molecule attached to the surface. The array can include a single copy of the indexed fragment at each site (also referred to as a feature), or multiple copies having the same sequence can be present at each site or feature. The multiple copies can be generated by amplification methods such as bridge amplification or emulsion PCR described in more detail herein.
[0176] The method described in this specification can use, for example, an array having features of any of various densities, including at least about 10 features / cm 2 , 100 features / cm 2 , 500 features / cm 2 , 1,000 features / cm 2 , 5,000 features / cm 2 , 10,000 features / cm 2 , 50,000 features / cm 2 , 100,000 features / cm 2 , 1,000,000 features / cm 2 , 5,000,000 features / cm 2 , or more.
[0177] An advantage of the method described in this specification is that for multiple cm 2is to provide rapid and efficient, parallel detection. Thus, the present disclosure provides an integrated system capable of preparing and detecting nucleic acids using techniques known in the art, such as those exemplified herein. Thus, the integrated system of the present disclosure can include fluid components capable of delivering amplification reagents and / or sequencing reagents to one or more immobilized indexed fragments, and the system can include components such as pumps, valves, reservoirs, fluid lines, etc. The flow cell can be configured and / or used in an integrated system for detecting target nucleic acids. Exemplary flow cells are described, for example, in U.S. Patent Application Publication No. 2010 / 0111768 and U.S. Patent Application No. 13 / 273,666. As exemplified for the flow cell, one or more of the fluid components of the integrated system can be used in amplification methods and detection methods. Taking an embodiment of nucleic acid sequencing as an example, one or more of the fluid components of the integrated system can be used for delivering sequencing reagents in the amplification methods described herein and sequencing methods as exemplified above. Alternatively, the integrated system can include separate fluid systems for performing amplification methods and detection methods. Examples of integrated sequencing systems capable of producing amplified nucleic acids and also determining the sequence of the nucleic acids include, but are not limited to, the MiSeq™ platform (Illumina, Inc., San Diego, CA), and the apparatus described in U.S. Patent Application No. 13 / 273,666.
[0178] Detection of Rare Events
[0179] The present disclosure also provides methods for identifying and / or characterizing rare events. Currently, methods for characterizing rare events within a population without enriching for the rare event are costly and difficult. When enrichment is used, the selection is typically based on some biological characteristics of the cells, such as size, morphology, or the presence or absence of distinguishable molecules such as proteins or glycans on the cell surface. This limits the types of events that can be identified. The methods presented herein represent a significant advancement in the ability to identify and / or characterize the presence of rare events. Generally, the invention provides for the identification, enrichment, and sequencing-based characterization of rare single-cell subsets present within libraries of millions or billions of cells. Identification of rare single cells can be used to create a cell database that researchers can use to determine cells for further analysis.
[0180] Examples of rare events include, but are not limited to, rare cells within a large population of cells. Types of rare cells include, but are not limited to, cell classes, species types, and disease states or risks. Examples of rare cell classes include, but are not limited to, cells from individuals having modifications in the genome, transcriptome, or epigenome. Examples of rare species types include, but are not limited to, prokaryotic cells, eukaryotic cells, or fungal cells. Examples of rare cells associated with disease states or risks include, but are not limited to, cancer cells.
[0181] Rare events are typically identified by biological features (usually the presence or absence of nucleotide sequences) that correlate with the rare event. In one embodiment, the biological feature is a biomolecule such as a protein, glycan, proteoglycan, or lipid. The biomolecule can be tagged with a nucleic acid conjugated to a compound, such as an antibody, that specifically binds to the biomolecule. The biological feature can be known in advance (e.g., known prior to the implementation of the method and referred to as predetermined) or newly discovered (e.g., the biological feature is identified after targeted sequencing or comprehensive sequencing as described herein).
[0182] Examples of biological features related to the genome include, but are not limited to, modifications in immune cells such as gene rearrangements. Examples of biological features related to the transcriptome include the expression of one or more specific genes or RNA molecules, or the expression of a specific protein. Examples of biological features related to the epigenome include, but are not limited to, epigenetic patterns such as methylation marks, methylation patterns, and accessible DNA, or the expression of specific proteins correlated with epigenetic changes. Examples of biological features correlated with the type of rare species include 16s rRNA or rDNA, 18s rRNA or rDNA, and internal transcribed spacer (ITS) rRNA / rDNA, or the expression of specific proteins by rare species. Examples of biological features related to disease states or risks include germline or somatic cells having mutant DNA sequences or expression patterns of RNA and / or proteins correlated with diseases such as cancer.
[0183] The method can include identifying members of a sequencing library (individual modified target nucleic acids) that include rare events. In one embodiment, the method can include screening a suspected sequencing library that includes rare events. Screening a sequencing library typically includes determining, for the sequences of two types of nucleotide regions present in the library, (i) biological features correlated with rare events, and (ii) indexes present in the members of the library. In one embodiment, the sequences of two or more biological features can be determined.
[0184] In one embodiment, the nucleotide sequence of the biological feature is identified by target sequencing. Target sequencing methods are known in the art and can include the use of primers that hybridize to approach the biological feature in terms of position and orientation serving as the starting site for sequencing. For example, when the biological feature is the presence or absence of a specific single nucleotide polymorphism (SNP), primers that specifically anneal to nucleotides close to the SNP can be designed. In another example, when the biological feature is a protein, primers that specifically anneal to the nucleotides of a nucleic acid attached to a compound specifically bound to a biomolecule can be designed. As a result, those skilled in the art can obtain sequence data that enables the identification of members of a library containing the target biological feature. Determining the sequence of the index present in the members of the sequencing library is an everyday part of single-cell combinatorial indexing methods.
[0185] Next, the sequence data from the targeted sequencing of the biological features and the sequencing of the index are analyzed using bioinformatics methods, which are conventional methods, to identify these combinations of index sequences that are present in the same library member as the biological feature. This correlation between the biological feature and the index sequence identifies a subset of the members of the library, and each member includes the unique classification of the biological feature and the index sequence, as well as the creation of a cell database. Each unique classification of the index sequence, also referred to herein as a "marker index sequence", is also present in other members of the library derived from the same cell or nucleus, for example, the index library of interest. In one embodiment, the marker index sequences are contiguous indices, i.e., a plurality of index sets that are present in the library members in rows having 0, 1, 2, 3, 4, or more nucleotides between each index. As described herein, these marker index sequences can be used to focus subsequent sequencing efforts on these members of the library derived from cells or nuclei having the biological feature, thus reducing costs.
[0186] The method can further include modifying the sequencing library to increase the representation of these members derived from cells or nuclei having the biological feature. Modifying can include enrichment (e.g., positive selection of these rare members of the library that contain the desired marker index sequence) or depletion (e.g., negative selection such as the selective removal of abundant members of the library that do not contain the desired marker index sequence).
[0187] Enrichment and depletion can involve using a marker index array. Methods for enrichment and depletion are known in the art and include, but are not limited to, marker index array-specific amplification (e.g., adapter-ligated PCR), hybrid capture, and hybridization-based methods such as CRISPR(d)Cas9. Enrichment methods and depletion methods benefit from using nucleotide sequences that specifically hybridize to a desired marker index array. Thus, enrichment or depletion can be performed on a set of multiple indexes (see FIG. 5B) present in library members in a continuous index, i.e., a row having 0, 1, 2, 3, 4 or more nucleotides between each index. A continuous index correlated with a desired biological feature can be reliably selected and retained, resulting in enrichment of the desired library members. Alternatively, a continuous index not correlated with a desired biological feature can be selected and removed, resulting in depletion of library members correlated with abundant cells and effectively enrichment of library members correlated with a desired biological feature. In one embodiment, enrichment can involve target amplification. For example, after construction of a sequencing library, an amplification reaction can be used to specifically amplify library members containing a target biological feature. In one embodiment, specific amplification can be achieved using a biological feature-specific primer designed to anneal to a nucleotide sequence having the biological feature and a second primer that anneals to one side of all members of the library. The biological feature-specific primer can include one or more indexes and / or universal sequences at its 5' end.
[0188] The full length of the contiguous index depends on the size of the probe required for specific hybridization between the probe and a member of the library having the desired marker index array. In some embodiments, the full length of the contiguous index (and thus the marker index array) is at least 40 nucleotides, at least 45 nucleotides, at least 50 nucleotides, or at least 55 nucleotides, and 80 nucleotides or less, 75 nucleotides or less, 70 nucleotides or less, or 65 nucleotides or less. In one embodiment, the full length of the contiguous index is 60 nucleotides.
[0189] By using either enrichment or depletion, a sub-library can be obtained that contains an increased representation of these members of the library derived from cells or nuclei having the biological feature. Comprehensive sequencing of the sub-library can be performed using conventional methods such as those described herein. Since the representation is increased sufficiently, comprehensive sequencing requires significantly fewer resources and is thus cost-effective. By using comprehensive sequencing of the sub-library, one or more additional biological features that were previously unknown can be identified.
[0190] Use
[0191] The methods provided by the present disclosure can be readily incorporated into essentially any application, including the preparation of sequencing libraries such as whole genomes, transcriptomes, epigenomes, accessible (e.g., ATAC), and three-dimensional conformational states (e.g., HiC). A number of sequencing library methods are known to those skilled in the art that can be used to construct whole genome or target libraries (see, e.g., "Sequencing Methods Review" available at genomics.umn.edu / downloads / sequencing-methods-review.pdf).
[0192] In these embodiments for detecting rare events, the methods provided by the present disclosure can be readily incorporated into essentially any application using single-cell combinatorial indexing (sci) methods, including but not limited to whole genome (e.g., sci-WGS-seq), epigenome (e.g., sci-MET-seq), accessible (e.g., sci-ATAC-seq), transcriptome (sci-RNA-seq), and three-dimensional structure (sci-HiC-seq). In some embodiments, the application involves using three-dimensional single-cell combinatorial indexing, including proximity ligation using a ligation long-read method with crosslinking. In some embodiments, the application is a co-assay that simultaneously evaluates two or more different analytes or information from a sample. Examples of analytes include, but are not limited to, DNA, RNA, and proteins (e.g., surface proteins). Examples include assays that analyze the whole genome and transcriptome, or ATAC and transcriptome (Ma et al., 2020, bioRxiv, DOI: doi.org / 10.1016 / j.cell.2020.09.056).
[0193] In some embodiments, the use is metagenomics (the study of genetic material recovered directly from environmental samples). Examples of environments include those associated with agriculture (e.g., soil), biofuels (e.g., the microbiota that convert biomass), biotechnology (e.g., the microbiota that produce bioactive compounds), and the gut microbiota (e.g., the microbiota present in the human or animal microbiome). The genetic material can be present in prokaryotic and / or eukaryotic microorganisms (both single-celled and multi-celled), such as fungal cells. The methods described herein can be used to identify rare cells, whether or not they can be cultured. Biological features that can be used to identify rare events in metagenomics include, but are not limited to, 16s rRNA or rDNA, 18s rRNA or rDNA, and internal transcribed spacer (ITS) rRNA / rDNA, or proteins encoded by the microorganism. After identification, the rare cells can be comprehensively sequenced.
[0194] In some embodiments, the present application relates to a disease state or risk. Rare events such as single nucleotide polymorphisms (SNPs) and / or biomarkers that correlate with a disease or risk of disease can be identified, and these cells having the SNPs and / or biomarkers are comprehensively sequenced. For example, a liquid biopsy of circulating cells in a subject's bloodstream, or a tissue biopsy of cells, can be analyzed for rare events related to a disease or risk of disease. Rare events that can be assayed include, but are not limited to, somatic driver mutations that enable the assignment of a particular cancer. A related use is to obtain samples from a subject over a period of time, select cancerous cells or nuclei, and then comprehensively sequence a subset of the tumor cells to fully characterize and track the progression of the tumor.
[0195] In some embodiments, the present application relates to immune cells. The immune cells undergo rearrangement of specific genes related to the acquired ability of the immune system to identify foreign molecules. Examples of immune cells that undergo gene rearrangement include, but are not limited to, T cells (e.g., rearrangement of the T cell receptor), antigen-presenting cells (e.g., rearrangement of genes encoding major histocompatibility complex proteins), and B cells (e.g., rearrangement of genes encoding antibodies). The biological characteristics related to the modification of immune cells can be, but are not limited to, specific rearrangements or proteins obtained from specific rearrangements. Immune cells with specific modifications, including but not limited to, repertoire characteristics and evolution of the T cell receptor, can be fully characterized and tracked. In another embodiment, the present application relates to cell differentiation. For example, differentiation events such as the correlation between accessibility and expression can be evaluated using expression levels and / or methylation in different regions.
[0196] Non-limiting exemplary embodiments of the present disclosure are shown in FIG. 6. In this embodiment, a method for identifying and characterizing a T cell receptor repertoire can include providing a plurality of cells (FIG. 6, block 600) and distributing a subset of the cells into a plurality of compartments (FIG. 6, block 601). The plurality of cells can be, for example, from a blood sample or a lymph node sample. The nucleic acids present in the cells of each compartment are modified by index insertion (FIG. 6, block 602), and then the cells are pooled (FIG. 6, block 603). Additional indices are added by a "split and pool" process that repeats the distribution (FIG. 6, block 601), index addition (FIG. 6, block 602), and subset pooling (FIG. 6, block 603). In one embodiment, each index is added to the same side of the members of the library to result in consecutive indices (see FIG. 5B). Optionally, a universal sequence can be added with one or more indices. After adding the last index, the library of nucleic acids in the nuclei or cells is pooled (FIG. 6, block 603) and further processed to enable identification of T cell receptors containing specific nucleotide sequences capable of binding to a biological feature of interest, such as a biomolecule of a microorganism or virus, and sequencing of the indices associated with the biological feature of interest for target sequencing of the biological feature (FIG. 6, block 604). Sequence analysis (FIG. 6, block 605) is used to identify marker index sequences, i.e., unique classifications of the index sequences. The identified marker index sequences are either (i) those that correlate with a biological feature and thus identify members of the library derived from rare cells, or (ii) those that do not correlate with a biological feature and thus identify members of the library derived from abundant cells. The subsequent steps of this exemplary embodiment describe depletion of abundant members of the library, but the method can be modified as described herein to include enrichment of rare library members.Specific oligonucleotide or guide RNA sequences can be designed to hybridize with marker index sequences that correlate with members of a library derived from abundant cells (Figure 6, block 606), and then, for example, by using hybridization capture or CRISPR digestion, a sequencing library of members derived from abundant cells can be depleted (Figure 6, 607). As a result, a modified library is obtained that contains an increased representation of these members derived from cells with biological characteristics. Members of the modified sequencing library can be subjected to comprehensive sequencing (Figure 6, block 608). Alternatively, the modified library can be subjected to further rounds of enrichment and / or depletion until the representation of the desired members of the library is sufficient to meet the criteria for characterization. For example, members of the modified library can be subjected to a second sequencing, the marker index is identified, and specific oligonucleotide or guide RNA sequences are designed and used to deplete or enrich the modified library.
[0197] In some embodiments, the use includes using a continuous index. A non-limiting exemplary embodiment of an approach for creating a sequencing library using a continuous index is shown in FIG. 7. After partitioning a subset of cells or nuclei, for example, by tagging, a first compartment-specific index I1 can be added to the DNA molecules 705 present in the cells or nuclei (FIG. 7, step 701). If the primary source of nucleic acid is RNA, the nucleic acid can be converted to DNA using methods such as cDNA synthesis prior to tagging. As a result, a library of modified nucleic acids present in the cells or nuclei is obtained, and each modified nucleic acid 706 contains the compartment-specific index I1 at each end. The subsets can be pooled, and the resulting ends of the modified target nucleic acids can be repaired, if necessary, for example, by 3' fill-in. In one embodiment, the 5' end of the modified target nucleic acid can be phosphorylated. In one embodiment, the next step of adding a second index can be facilitated by adding an overhang (e.g., G, C, or polyA tail) to the 3' end of the modified target nucleic acid. The pooled cells or nuclei are partitioned into a second set of compartments, and a second compartment-specific index I2 can be added, for example, by ligation of an adapter having a suitably modified 3' end, e.g., a T-tailed 3' end (FIG. 7, step 702). Thereby, cells or nuclei containing a library of modified nucleic acids are obtained, and each modified nucleic acid 707 contains two compartment-specific indexes I1 and I2 at each end. The ends of the modified target nucleic acids can be modified, for example, by 5' phosphorylation and / or modification of the 3' end with a polyA tail, or addition of G or C to the 3', to facilitate the addition of the next index. Optionally, pooling and addition of another compartment-specific index can be repeated to add an appropriate number of indexes. In one embodiment, when adding a final compartment-specific index I3 to the partitioned cells or subset, an adapter having a universal sequence can be included (FIG. 7, step 703). For example, a mismatched adapter can be added to each end to obtain a modified nucleic acid 708. Examples of universal sequences include those used to immobilize library members on an array (P5 and P7).The mismatch adapter can also contain universal sequences useful for sequencing, or in some embodiments, can amplify the modified nucleic acid 708 (Figure 7, step 704), and add universal sequences (i5 and i7) useful for sequencing to obtain the modified nucleic acid 709. The modified nucleic acid 709 can be used in target sequencing to identify marker index sequences that correlate with biological features useful for subsequent enrichment and / or deletion.
[0198] A non-limiting exemplary embodiment combining enrichment with target amplification is shown in Figure 8. In this embodiment, a single-cell combinatorial library has been created (e.g., Figure 3, block 35; Figure 4, block 47; Figure 6, block 605), and the resulting modified nucleic acid (e.g., modified nucleic acid 709 in Figure 7) is subjected to an amplification reaction that specifically amplifies library members containing the biological feature of interest. The modified nucleic acid 802 with consecutive indices can contact a primer 803 that may include two domains, namely, a 3' domain designed to anneal to the nucleotide sequence with the biological feature, and a 5' domain having one or more universal sequences or their complements, e.g., i7 and P7. The amplification reaction includes a second primer 804 that anneals to one side of all members of the library. Amplification 801 results in a modified nucleic acid 805 having a partition-specific index I1-3 at one end and a universal sequence added together with a two-domain primer targeting the biological feature at the other end. The amplified modified target nucleic acid can be used in target sequencing and sequencing to identify marker index sequences that correlate with the biological feature of interest.
[0199] Also provided herein is a kit. In one embodiment, the kit is for preparing a sequencing library. In one embodiment, the kit includes one transpososome complex and includes a transposon recognition site such that a universal sequence can be inserted into a target nucleic acid. In another embodiment, the kit includes two transpososome complexes, each complex including a transposon recognition site having a different universal sequence such that a universal sequence can be inserted into a target nucleic acid. In another embodiment, the kit includes components for adding at least one, two, or three indexes to a nucleic acid. The kit may also include other components useful for making a sequencing library. For example, the kit may include at least one enzyme that mediates ligation, primer extension, or amplification to process a DNA molecule to include an index. The kit may include a nucleic acid having an index sequence.
[0200] The components of the kit are typically in a suitable packaging material in an amount sufficient for at least one assay or use. Optionally, other components such as buffers and solutions may be included. Typically, instructions for use of the packaged components are also included. As used herein, the phrase "packaging material" refers to one or more physical structures used to contain the contents of the kit. The packaging material is generally constructed by conventional methods to provide a sterile, contaminant-free environment. The packaging material may have a label indicating that the components can be used to make a sequencing library. In addition, the packaging material includes instructions describing how to use the materials in the kit. As used herein, the term "package" refers to a container such as glass, plastic, paper, foil, etc. that can hold the components of the kit within certain limits. "Instructions for use" typically include specific expressions that describe at least one assay parameter such as reagent concentration, or relative amounts of reagents and samples to be mixed, duration of the reagent / sample mixture, temperature, buffer conditions, etc.
[0201] Composition
[0202] During or after the generation of a sequencing library, a large number of molecules and compositions may be obtained. For example, molecules or compositions that may result include modified target nucleic acids flanked on one or both sides by a contiguous index. The contiguous index may include 1, 2, 3, 4, 5, 6, or more indices in a row, and each index is separated from other indices by 1, 2, 3, 4, or more nucleotides. In some embodiments, the total length of the contiguous index is at least 40 nucleotides, at least 45 nucleotides, at least 50 nucleotides, or at least 55 nucleotides and 80 nucleotides or less, 75 nucleotides or less, 70 nucleotides or less, or 65 nucleotides or less. A library or composition containing a plurality of such modified target nucleic acids may be obtained. Pooled libraries and compositions containing pooled libraries of such polynucleotides may be obtained.
[0203] Exemplary embodiments
[0204] Embodiment 1. A method for identifying a subpopulation of cells comprising a biological characteristic, comprising: (a) providing a single cell sequencing library, the sequencing library comprising a plurality of modified target nucleic acids, the modified target nucleic acids comprising at least one index sequence, and (b) interrogating the sequencing library by targeted sequencing to identify the index sequence present in the modified target nucleic acid that is the same as the biological characteristic, the index sequence associated with the biological characteristic being a marker index sequence, and (c) modifying the sequencing library to obtain a sublibrary, The sub-library contains an increased representation of the modified target nucleic acid containing the marker index array as compared to other modified target nucleic acids present within the sequencing library that do not contain the marker index array, and (d) determining the nucleotide sequence of the modified target nucleic acid containing the marker index array, a method comprising:
[0205] Embodiment 2. The method according to embodiment 1, wherein the single cell sequencing library contains nucleic acids from a plurality of samples.
[0206] Embodiment 3. The method according to any one of embodiments 1 to 2, wherein the plurality of samples includes (i) samples of the same tissue obtained from different organisms, (ii) samples of different tissues from one organism, or (iii) samples of different tissues from different organisms.
[0207] Embodiment 4. The method according to any one of embodiments 1 to 3, wherein in step (b), two or more marker index arrays are identified.
[0208] Embodiment 5. The method according to any one of embodiments 1 to 4, wherein the single cell combinatorial sequencing library contains target nucleic acids representing the whole genome or a genomic subset of cells or nuclei.
[0209] Embodiment 6. The method according to any one of embodiments 1 to 5, wherein the genomic subset contains target nucleic acids representing the transcriptome, accessible chromatin, DNA, three-dimensional structural state, or proteins of cells or nuclei.
[0210] Embodiment 7. The method according to any one of embodiments 1 to 6, wherein modifying includes enriching the modified target nucleic acid containing the marker index array.
[0211] Embodiment 8. The method according to any one of embodiments 1 to 7, wherein the enrichment includes a hybridization-based method.
[0212] Embodiment 9. The hybridization-based method is the method according to any one of Embodiments 1 to 8, including hybrid capture, amplification, or CRISPR(d)Cas9.
[0213] Embodiment 10. Modifying includes depleting the modified target nucleic acid without the marker index sequence, and is the method according to any one of Embodiments 1 to 9.
[0214] Embodiment 11. The depletion includes a hybridization-based method, and is the method according to any one of Embodiments 1 to 10.
[0215] Embodiment 12. The hybridization-based method is the method according to any one of Embodiments 1 to 11, including hybrid capture, amplification, or CRISPR(d)Cas9.
[0216] Embodiment 13. The biological feature includes a nucleotide sequence indicating the type of species, and is the method according to any one of Embodiments 1 to 12.
[0217] Embodiment 14. The type of species includes cell species, and is the method according to any one of Embodiments 1 to 13.
[0218] Embodiment 15. The biological feature includes nucleotides of the 16s subunit, 18s subunit, or ITS non-transcribed region, and is the method according to any one of Embodiments 1 to 14.
[0219] Embodiment 16. The biological feature includes a nucleotide sequence indicating the cell class, and is the method according to any one of Embodiments 1 to 15.
[0220] Embodiment 17. The cell class includes an expression pattern, an epigenetic pattern, immunogenetic recombination, or a combination thereof, and is the method according to any one of Embodiments 1 to 16.
[0221] Embodiment 18. The epigenetic pattern is the method according to any one of Embodiments 1 to 17, including a methylation label, a methylation pattern, accessible DNA, or a combination thereof.
[0222] Embodiment 19. The biological characteristic is the method according to any one of Embodiments 1 to 18, including a nucleotide sequence indicating a disease state or risk.
[0223] Embodiment 20. The disease state or risk is the method according to any one of Embodiments 1 to 19, including a mutant DNA sequence, a mutant expression pattern, or a mutant epigenetic pattern correlated with a disease.
[0224] Embodiment 21. The mutant DNA sequence includes at least one single nucleotide polymorphism, and is the method according to any one of Embodiments 1 to 20.
[0225] Embodiment 22. The mutant expression pattern includes the expression of a biomarker, and is the method according to any one of Embodiments 1 to 21.
[0226] Embodiment 23. The mutant epigenetic pattern includes a methylation label and a methylation pattern, and is the method according to any one of Embodiments 1 to 22.
[0227] Embodiment 24. The modified target nucleic acid includes consecutive indexes of at least two compartment-specific index sequences, and there are no more than 7 nucleotides between the two index sequences, and is the method according to any one of Embodiments 1 to 23.
[0228] Embodiment 25. The consecutive indexes are present at each end of the modified target nucleic acid, and is the method according to any one of Embodiments 1 to 24.
[0229] Embodiment 26. The length of the consecutive indexes is at least 55 nucleotides, and is the method according to any one of Embodiments 1 to 25.
[0230] Embodiment 27. One copy of the continuous index is the method according to any one of Embodiments 1 to 26, which is present in the modified target nucleic acid.
[0231] Embodiment 28. Two copies of the continuous index are the method according to any one of Embodiments 1 to 27, which is present in the modified target nucleic acid.
[0232] Embodiment 29. The plurality of modified target nucleic acids of the sequencing library are the method according to any one of Embodiments 1 to 28, which represent at least 100,000 different cells or nuclei.
[0233] Embodiment 30. Providing a single-cell combinatorial sequencing library includes processing a sample to create the library, wherein the sample is a metagenomic sample obtained from an organism, the method according to any one of Embodiments 1 to 29.
[0234] Embodiment 31. The organism is a mammal, the method according to any one of Embodiments 1 to 30.
[0235] Embodiment 32. The metagenomic sample includes a tissue suspected of containing symbiotic or pathogenic microorganisms, the method according to any one of Embodiments 1 to 31.
[0236] Embodiment 33. The microorganism is a prokaryote or a eukaryote, the method according to any one of Embodiments 1 to 32.
[0237] Embodiment 34. The metagenomic sample includes a microbiome sample, the method according to any one of Embodiments 1 to 33.
[0238] Embodiment 35. Providing a single-cell combinatorial sequencing library includes processing a sample to create the library, wherein the sample is from an organism, the method according to any one of Embodiments 1 to 34.
[0239] Embodiment 36. The method according to any one of Embodiments 1 to 35, wherein the organism is a mammal.
[0240] Embodiment 37. The method according to any one of Embodiments 1 to 36, wherein the primary source of nucleic acid from the sample contains RNA.
[0241] Embodiment 38. The method according to any one of Embodiments 1 to 37, wherein the RNA contains mRNA.
[0242] Embodiment 39. The method according to any one of Embodiments 1 to 38, wherein the primary source of nucleic acid from the sample contains DNA.
[0243] Embodiment 40. The method according to any one of Embodiments 1 to 39, wherein the DNA contains whole-cell genomic DNA.
[0244] Embodiment 41. The method according to any one of Embodiments 1 to 40, wherein the whole-cell genomic DNA contains nucleosomes.
[0245] Embodiment 42. The method according to any one of Embodiments 1 to 41, wherein the primary source of nucleic acid from the sample contains cell-free DNA.
[0246] Embodiment 43. The method according to any one of Embodiments 1 to 42, wherein the sample contains cancer cells.
[0247] Embodiment 44. Providing a single-cell combinatorial sequencing library includes preparing the library using a single-cell combinatorial indexing method selected from single-nucleus transcriptome sequencing, single-cell transcriptome sequencing, single-cell transcriptome and transposon-accessible chromatin sequencing, single-nucleus whole-genome sequencing, single-nucleus sequencing of transposon-accessible chromatin, single-cell epitope sequencing, sci-HiC, and sci-MET, according to any one of Embodiments 1 to 43.
[0248] Embodiment 45. The method according to any one of Embodiments 1 to 44, which includes providing two different single-cell combinatorial sequencing libraries from each cell or nucleus.
[0249] Embodiment 46. The two different single-cell combinatorial sequencing libraries are selected from single-cell combinatorial indexing methods including single-nucleus transcriptome sequencing, single-cell transcriptome sequencing, single-cell transcriptome and transposon-accessible chromatin sequencing, single-nucleus whole-genome sequencing, single-nucleus sequencing of transposon-accessible chromatin, sci-HiC, and sci-MET, according to any one of Embodiments 1 to 45.
[0250] Embodiment 47. The method according to any one of Embodiments 1 to 46, further including performing a sequencing procedure to determine the nucleotide sequence of the nucleic acid.
[0251] Embodiment 48. A method for preparing a sequencing library containing nucleic acids from a plurality of single nuclei or single cells, comprising: (a) providing a plurality of nuclei or cells, wherein the nuclei or cells contain nucleosomes; (b) contacting the plurality of nuclei or cells with a transpososome complex containing a transposase and a universal sequence, the contacting further including conditions suitable for integrating the universal sequence into the DNA nucleic acid to result in a double-stranded DNA nucleic acid containing the universal sequence; (d) distributing the plurality of nuclei or cells into a first plurality of compartments, wherein each compartment contains a subset of the nuclei or cells; (e) processing the DNA molecules within each subset of the nuclei or cells to generate indexed nuclei or cells. The process involves adding a first compartment-specific index sequence to the DNA nucleic acid present in each subset of nuclei or cells, resulting in indexed nucleic acids present in the indexed nuclei or cells. The process includes ligation, primer extension, hybridization, amplification, or combinations thereof. (g) combining the indexed nuclei or cells to generate pooled indexed nuclei or cells. A method comprising this.
[0252] Embodiment 49. The method according to claim 48, wherein providing comprises providing a plurality of nuclei or cells within a plurality of compartments, each compartment containing a subset of the nuclei or cells, and contacting comprises contacting each compartment with a transpososome complex, and the method further comprises combining the nuclei or cells after contacting to generate pooled nuclei or cells.
[0253] Embodiment 50. The method according to any one of embodiments 48 to 49, wherein providing comprises subjecting the nuclei to a chemical treatment to generate nucleosome-depleted nuclei while maintaining the integrity of the isolated nuclei.
[0254] Embodiment 51. Distributing the pooled indexed nuclei or cells containing the indexed nuclei or cells into a second plurality of compartments, where each compartment contains a subset of the nuclei or cells, processing the DNA molecules within each subset of the nuclei or cells to generate doubly indexed nuclei or cells, where the processing involves adding a second compartment-specific index sequence to the DNA nucleic acid present in each subset of the nuclei or cells, resulting in doubly indexed nucleic acids present in the indexed nuclei or cells, the processing includes ligation, primer extension, hybridization, amplification, or combinations thereof. The method according to any one of embodiments 48 to 50, further comprising combining the nuclei or cells with dual indices to generate pooled nuclei or cells with dual indices.
[0255] Embodiment 52. Distributing the pooled nuclei or cells containing nuclei or cells with dual indices into a third plurality of compartments, wherein each compartment contains a subset of the nuclei or cells, processing the DNA molecules within each subset of the nuclei or cells to generate nuclei or cells with triple indices, wherein processing comprises adding a third compartment-specific index sequence to the DNA nucleic acids present in each subset of the nuclei or cells, resulting in nucleic acids with triple indices present in the indexed nuclei or cells, wherein processing comprises ligation, primer extension, hybridization, amplification, or a combination thereof, The method according to any one of embodiments 48 to 51, further comprising combining the nuclei or cells with triple indices to generate pooled nuclei or cells with triple indices.
[0256] Embodiment 53. The method according to any one of embodiments 48 to 52, wherein the step of distributing comprises dilution.
[0257] Embodiment 54. The method according to any one of embodiments 48 to 53, wherein the compartments comprise wells, microfluidic compartments, or droplets.
[0258] Embodiment 55. The method according to any one of embodiments 48 to 54, wherein the compartments of the first plurality of compartments contain 50 to 100,000,000 nuclei or cells.
[0259] Embodiment 56. The method according to any one of embodiments 48 to 55, wherein the compartments of the second plurality of compartments contain 50 to 100,000,000 nuclei or cells.
[0260] Embodiment 57. The compartments of the third plurality of compartments contain 50 to 100,000,000 nuclei or cells, and the method according to any one of Embodiments 48 to 56.
[0261] Embodiment 58. Contacting comprises contacting each subset with two transpososome complexes, one transpososome complex comprising a first transposase comprising a first universal sequence, and the second transpososome complex comprising a second transposase comprising a second universal sequence, and contacting further comprises conditions suitable for incorporating the first universal sequence and the second universal sequence into a DNA nucleic acid to yield a double-stranded DNA nucleic acid comprising the first universal sequence and the second universal sequence, and the method according to any one of Embodiments 48 to 57.
[0262] Embodiment 59. Adding a compartment-specific index sequence comprises a two-step process of adding a nucleotide sequence comprising a universal sequence to a nucleic acid and then adding a compartment-specific index sequence to the nucleic acid, and the method according to any one of Embodiments 48 to 58.
[0263] Embodiment 60. Further comprising obtaining an indexed nucleic acid from the pooled indexed nuclei or cells, thereby further comprising creating a sequencing library from a plurality of nuclei or cells, and the method according to any one of Embodiments 48 to 59.
[0264] Embodiment 61. Further comprising obtaining a double-indexed nucleic acid from the pooled double-indexed nuclei or cells, thereby further comprising creating a sequencing library from a plurality of nuclei or cells, and the method according to any one of Embodiments 48 to 60.
[0265] Embodiment 62. Further comprising obtaining a triple-indexed nucleic acid from the pooled triple-indexed nuclei or cells, thereby further comprising creating a sequencing library from a plurality of nuclei or cells, and the method according to any one of Embodiments 48 to 61.
[0266] Embodiment 63. further comprising the step of providing a surface comprising a plurality of amplification sites, the amplification sites comprising at least two populations of ligated single-stranded capture oligonucleotides having free 3' ends, contacting the surface comprising the amplification sites with a nucleic acid fragment comprising one, two, or three index sequences under conditions suitable to generate a clonal population of amplicons from each of a plurality of amplification sites, each comprising a clonal population of amplicons from an individual fragment comprising a plurality of indexes; the method according to any one of embodiments 48 to 62.
[0267] Embodiment 64. A method for preparing a nucleic acid library, comprising: (a) providing a plurality of samples, each sample comprising a plurality of cells or nuclei, the plurality of cells or nuclei in each sample being present in one or more separate compartments; (b) contacting the plurality of nuclei or cells with a transpososome complex comprising a transposase and a universal sequence, under conditions such that the transpososome complex does not comprise an index sequence, the contacting further comprising conditions suitable to incorporate the universal sequence into the nucleic acid; (c) adding a first index sequence to the nucleic acid of each separate compartment; (d) combining the cells or nuclei of the separate compartments; (e) distributing the cells or nuclei into a plurality of compartments; (f) adding a second index sequence to the nucleic acid of the plurality of compartments.
[0268] Embodiment 65. The method according to embodiment 64, wherein the first index sequence, the second index sequence, or a combination thereof is added by ligation, primer extension, hybridization, amplification, or a combination thereof.
[0269] Embodiment 66. The method according to any one of Embodiments 64 to 65, wherein steps (d) to (e) are repeated to add a third or more index array to cells or nuclei in a plurality of compartments.
[0270] Embodiment 67. The method according to any one of Embodiments 64 to 66, wherein the plurality of nuclei or cells are fixed.
[0271] Embodiment 68. The method according to any one of Embodiments 64 to 67, further comprising amplifying the indexed nucleic acid after step (c) or step (f).
[0272] Embodiment 69. The method according to any one of Embodiments 64 to 68, further comprising combining nucleic acids in a plurality of compartments and determining the sequence of the nucleic acids in step (g).
[0273] Embodiment 70. The method according to any one of Embodiments 64 to 69, further comprising performing a sequencing procedure to determine the nucleotide sequence of the nucleic acid.
[0274] Embodiment 71. A method for sequencing a single cell or a single nucleus, comprising: (a) uniquely indexing the nucleic acid of each cell or nucleus in a sample, thereby creating an indexed library of each cell or nucleus; (b) using a biological characteristic to identify one or more indexed libraries from step (a) that are of interest; (c) concentrating the indexed library of interest from step (b), thereby creating a concentrated library; (d) sequencing the concentrated library from step (c).
[0275] Embodiment 72. The method according to Embodiment 71, wherein the library is derived from DNA, RNA, or protein of a cell or nucleus.
[0276] Embodiment 73. The method according to any one of Embodiments 64 to 72, wherein the biological feature is DNA, RNA, or protein, or a combination thereof.
[0277] Embodiment 74. The method according to any one of Embodiments 64 to 73, wherein uniquely indexing in step (a) includes associating at least two different indexes with the nucleic acid of a cell or nucleus.
[0278] Embodiment 75. The method according to any one of Embodiments 64 to 74, wherein the at least two different indexes are consecutive indexes.
[0279] Embodiment 76. The method according to any one of Embodiments 64 to 75, wherein the enriched library is produced by positive enrichment.
[0280] Embodiment 77. The method according to any one of Embodiments 64 to 76, wherein positive enrichment includes amplification.
[0281] Embodiment 78. The method according to any one of Embodiments 64 to 77, wherein positive enrichment includes a capture agent.
[0282] Embodiment 79. The method according to any one of Embodiments 64 to 78, wherein positive enrichment includes a solid support.
[0283] Embodiment 80. The method according to any one of Embodiments 64 to 79, wherein the enriched library is produced by negative enrichment.
[0284] Embodiment 81. The method according to any one of Embodiments 64 to 80, wherein identifying the target indexed library in step (c) includes sequencing the index.
[0285] Embodiment 82. A method for sequencing a single cell or a single nucleus, comprising: (a) providing a sample, the sample comprising a plurality of nuclei or cells, (b) Associating a first index with each nucleus or cell in the sample; (c) Dividing the sample into a plurality of compartments; (d) Associating a second index with each nucleus or cell in the plurality of compartments; (e) Pooling the plurality of compartments; (f) Sequencing the pooled compartments; (g) Identifying combinations of the first index and the second index associated with biological features; (h) Using the identified combinations of the first index and the second index from step (g) to enrich biological features from the pooled compartments. A method comprising the steps.
[0286] Embodiment 83. A kit comprising: (a) A plurality of transpososome complexes, each transpososome complex comprising a transposase and a transposon sequence, the transposon sequence being unindexed; a plurality of transpososome complexes; (b) A first plurality of index oligonucleotides, the first plurality of index oligonucleotides comprising oligonucleotides having at least two different sequences; a first plurality of index oligonucleotides; (c) A ligase enzyme for use with the index oligonucleotides. A kit comprising the components.
[0287] Embodiment 84. The kit according to embodiment 83, further comprising a second plurality of index oligonucleotides, the second plurality of index oligonucleotides comprising oligonucleotides having a sequence different from that of the first plurality of index oligonucleotides.
[0288] Embodiment 85. The kit according to Embodiment 83 or 84, further comprising a third plurality of index oligonucleotides, wherein the third plurality of index oligonucleotides comprise oligonucleotides having a sequence different from those of the first plurality of index oligonucleotides and the second plurality of index oligonucleotides.
[0289]
Examples
[0290] The present disclosure is illustrated by the following examples. It should be understood that the specific examples, materials, amounts, and procedures should be broadly interpreted in accordance with the scope and spirit of the present disclosure described herein.
[0291] Example 1
[0292] Human Cell Atlas of Chromatin Accessibility during Development
[0293] Summary
[0294] The chromatin landscape of the human genome shapes cell-type-specific programs of gene expression. The inventors developed an improved assay for single-cell profiling of chromatin accessibility based on three levels of combinatorial indexing (sci-ATAC-seq3), applied it to 59 fetal samples representing 15 organs, and profiled nearly a million single cells in total. The inventors utilized cell types defined by gene expression in the same organ to annotate these data, constructed a catalog of hundreds of thousands of cell-type-specific DNA regulatory elements, investigated lineage-specific transcription factor properties, and cell-type-specific enrichment of complex traits inheritance. Together with the accompanying Human Cell Atlas of gene expression during development, these data contain a rich resource for exploring human biology.
[0295] Main Text
[0296] In recent years, single-cell methods, experiments, and atlases have been increasing rapidly. However, the overwhelming majority of those efforts have remained focused on single-cell gene expression, which reflects only one aspect of cell biology, developmental biology, and organismal biology. Other aspects, such as chromatin landscapes that shape gene expression programs, are equally important for investigations at single-cell resolution but suffer from the problem of relatively few scalable methods.
[0297] The single-cell combinatorial indexing (“sci”) framework involves splitting cells or nuclei and pooling them into wells, where molecular barcodes are introduced in situ each time to the species of interest (e.g., RNA or chromatin). Through successive in situ molecular barcoding, species within the same cell are labeled in correspondence with unique combinations of barcodes, and sci-assays for profiling chromatin accessibility (sci-ATAC-seq), gene expression (sci-RNA-seq), nuclear structure, genomic sequence, methylation, histone labeling, and other phenomena, as well as sci-co-assays for profiling, for example, chromatin accessibility and gene expression together (“CoBatch,” “Split-seq,” “Pagaired-seq,” and “dscATAC-seq” are methods that also rely on single-cell combinatorial indexing) have been developed.
[0298] Previously, chromatin accessibility in ~100,000 mammalian cells could be profiled via two-level sci-ATAC-seq, but the assay has several limitations. For example, it requires custom loading of the Tn5 enzyme with barcoded adapters, and 10 4 ~10 5is limited to cells that receive the same combination of barcodes, i.e., cells. To address these issues, the inventors developed an improved assay for single-cell profiling of chromatin accessibility based on three-level combinatorial indexing (sci-ATAC-seq3). In contrast to previous iterations of sci-ATAC-seq, this assay does not rely on a molecular barcode-tagged Tn5 complex (Figures 9; 10). Rather, the first two rounds of indexing are achieved by ligation to either end of a conventional, uniformly loaded Tn5 transposase complex (standard "Nextera"), and the final round of indexing is still via PCR. Compared to two-level sci-ATAC-seq, but similar to sci-RNA-seq3, sci-ATAC-seq3 significantly reduces the library preparation cost per cell as well as the collision rate. The theoretical collision rates for two-level indexing (96x384 wells) and three-level indexing (384x384x384 wells) are 12% and 1.3%, respectively, and the collision rate observed for a three-level "species mixing" experiment using equal numbers of pooled GM12878 cells and CH12.LX cells was estimated to be 4.0%, 10 6 opened the way for cell-scale experiments. This protocol no longer requires cell sorting. The inventors also optimized the selection of ligase and polymerase, kinase concentration, and oligo design and concentration to maximize the number of fragments recovered from each cell. Note that an explicit choice was made to maximize complexity at the expense of the specificity of accessible sites while maintaining enrichment within accessible regions. Picard was used to calculate the estimated total unique reads ("complexity") per cell and the Fraction of Reads in Transcription Start Site ("FRiTSS") per cell. Reads within 500 bp of the Gencode TSS were considered to be within the TSS. Specifically, it was found that by adjusting the fixed conditions, the sensitivity (i.e., complexity) and specificity (i.e., enrichment at accessible sites) of the assay could be adjusted.
[0299] Towards a human cell atlas of chromatin accessibility, sci-ATAC-seq3 was applied to 59 fetal samples representing 15 organs (adrenal gland, 2 regions of the cerebellum, eye, heart, intestine, kidney, liver, lung, muscle, pancreas, placenta, spleen, stomach, and thymus), profiling chromatin accessibility in all 1.6 million cells (Figures 1D - E). In Example 2, profiling of gene expression in 4 - 5 million cells from the same organs will be described based on overlapping sets of samples. The organs profiled span diverse systems and the most notable absences are bone marrow, bone, gonads, and skin.
[0300] Rapid and uniform processing of heterogeneous fetal tissues presents a difficult challenge. The inventors developed a new method for directly extracting nuclei from cryopreserved tissues that functions well across diverse tissue types and generates homogenates suitable for both sci-ATAC-seq3 and sci-RNA-seq3. Briefly, rapidly frozen tissue sections are wrapped in aluminum foil and pulverized into powder on dry ice using a cooled hammer. The tissue powder is then aliquoted, one for sci-ATAC-seq3 and the other for sci-RNA-seq3.
[0301] For sci-ATAC-seq3, samples were obtained from 23 fetuses with estimated gestational ages in the range of 89 - 125 days. Cells were lysed and nuclei were isolated using the published ATAC-seq cell lysis buffer, and the nuclei were fixed with formaldehyde prior to snap-freezing for future processing. For nuclei from each tissue, approximately 50,000 fixed nuclei were deposited across 4 wells of a 96-well plate and processed for tagging. After tagging, the first index, which also identified the tissue sample, was introduced by ligation to one of the free ends of the asymmetrically inserted transposase complex. After pooling and splitting, the second index was introduced by ligation to the other free end of the transposase complex. Following another round of pooling and splitting, the final index was added by PCR and the resulting amplicons were pooled for sequencing.
[0302] Sequenced the sci-ATAC-seq3 library of the third experiment from 5 Illumina NovaSeq experiments, generating over 50 billion reads in total. As a first QC check, the data was examined at the tissue level, i.e., before splitting into single cells. Downloaded all available single-end DNase-seq samples from fetal tissues from the ENCODE data portal and remapped them. Next, identified accessibility peaks in each "pseudo-bulk" sample and each ENCODE sample, merged them, and scored each sample for accessibility at each peak in the master list. However, the sci-ATAC-seq3 data was not very enriched at the peaks (central reads of the peak: 29% for ATAC-SEQ3 and 35% for ENCODE DNase-seq), but samples from the same tissue were similarly correlated for the two assays (central Spearman correlation: 0.93 for two samples from the same tissue in sci-ATAC-seq3 and 0.91 for DNase-seq), and sci-ATAC-seq3 had higher technical reproducibility (central Spearman correlation: 0.95). Furthermore, the samples were clustered into their respective tissues based on whether these aggregated profiles were analyzed with sci-ATAC-seq3 samples alone or by combining sci-ATAC-seq3 samples and DNase-seq samples using pairwise Spearman correlations for clustered samples.
[0303] Reads were split based on cell barcodes, and dynamic thresholds were applied as described above to identify 1,568,018 cells. A collision rate of ~5% was estimated for each of the three experiments from the chicken controls. Uniform Manifold Approximation and Projection (UMAP) visualization of cells corresponding to human sentinel tissue did not reveal obvious experimental batch effects. Considering the poor nucleosome banding of their fragment size distributions, three samples were dropped, and two more samples were dropped because few cells were captured. It is estimated that 91% - 99% of the centers of all unique fragments per cell were sequenced in these sci-ATAC-seq3 libraries for each tissue type.
[0304] After identifying accessibility peaks for each tissue, these were merged to generate a master set of 1.05 million sites. After scoring each cell for the presence or absence of reads at each site, low-quality cells were filtered out based on the total number of unique reads (sample-specific minimum in the range of 1,000 - 3,586), the proportion of reads overlapping the master set of accessible sites (sample-specific minimum in the range of 0.2 - 0.4), the proportion of reads falling near the TSS (+ / -1 kb; sample-specific minimum in the range of 0.05 - 0.15), and the doublet score obtained by adapting the Scrublet doublet detection algorithm originally developed for scRNA-seq data (excluding ~10% of cells with the highest doublet scores).
[0305] After these procedures, 790,957 single-cell chromatin accessibility profiles from 54 fetal samples remained. The total number of high-quality cells per tissue ranged from 2,421 (spleen) to 211,450 (liver). The median number of unique fragments per cell in this set was 6,042, the median of those overlapping the master set of accessible sites was 0.49, and the proportion falling near the TSS (+ / -1 kb) was 0.19.
[0306] The inventors used logarithmically transformed term frequency components to subject high-precision cells to latent semantic indexing (LSI) for each tissue. Although no obvious evidence of batch effects for different samples corresponding to the same tissue was observed, the Harmony algorithm was applied to align the samples in the PCA space for each tissue as a conservative means. Using the PCA space aligned for each tissue, Louvain clustering was then applied, initially obtaining 172 clusters across all tissues. UMAP was used to further reduce the dimensions of each tissue dataset.
[0307] Annotation of cell types
[0308] As shown by the inventors and others, the annotation of cell types within scATAC-seq datasets can be greatly simplified by leveraging scRNA-seq datasets. To partially automate the annotation of cell types for scATAC-seq data, first, the cell types within the scRNA-seq data for the same tissue were annotated as described in the manual. Second, gene-level accessibility scores were calculated for the scATAC-seq data, and the number of translocation events falling within the gene body extended by 2 kb upstream of their TSS was tabulated. Third, gene-cell matrices were used for each data type as input to an approach for finding possible correspondences between scRNA-seq clusters and scATAC-seq clusters based on non-negative least squares (NNLS) regression, thereby obtaining an initial "lift-over" set of automated annotations for scATAC-seq clusters. Finally, all automated annotations were manually reviewed by examining the pile-ups around marker genes for each cell type within each tissue and the assigned labels were corrected if deemed necessary. First, cell types were annotated with sci-RNA-seq data collected in the matching tissue based on marker gene expression. Louvain clusters were identified with the ATAC data for each tissue. Next, gene-level accessibility scores were calculated for each of these clusters, matched to the RNA clusters based on non-negative least squares (NNLS) regression, and in some cases, merging of Louvain clusters occurred. These first-round automated annotations were further refined by manually reviewing the cluster-specific accessibility landscapes around marker genes. The annotated cell types showed specific accessibility around the TSS of known marker genes. For each cell type or unannotated cluster, the accessibility near the TSS of known marker genes was summed and the scale was normalized to account for the difference in total reads per cell and the number of cells across the cell types.The data suggested that some of the unannotated clusters may not represent novel cell types, but rather technical artifacts (e.g., doublets). The inventors noted that while other approaches have shown great promise for multimodal integration of single-cell data, the cluster-to-cluster NNLS method was sufficient for the purposes herein and was found to be much less computationally intensive.
[0309] In total, 150 out of 172 (87%) clusters, or 163 out of 172 (95%) when including low-confidence labels, could be annotated. Some clusters received the same annotation within the same tissue and were thus merged, resulting in 124 annotations across all tissues. Some of these annotations were present across multiple tissues (e.g., erythroblasts in 4 tissues). By discarding across tissues, 54 (or 59 when including low-confidence labels and a 1:2 mapping) unique cell type annotations were obtained that mapped 1:1 to the annotations performed on the scRNA-seq dataset. Many of the scRNA-seq cell types not found in the chromatin accessibility data at this level of resolution are likely small clusters that were not sufficiently sampled to be detectable due to the low number of cells profiled in this study (~4M (RNA) vs ~800K (ATAC) high-quality cells). On the other hand, most of the 9 scATAC-seq clusters that remained completely unannotated were thought to be due to unfiltered doublets. This is because in the UMAP representation, they are characterized by accessibility in marker genes of multiple adjacent cell types.
[0310] Identification of lineage-specific TFs
[0311] Next, we sought to integrate and compare chromatin accessibility across all 15 organs in terms of cell type. To mitigate the effect of net differences in cell number per organ and / or cell type, 800 cells per cell type were randomly sampled per organ (or, if fewer than 800 cells of a given cell type were represented in a given organ, all cells were obtained), and UMAP visualization was performed. Reassuringly, cell types were shown across multiple organs clustered together, such as stromal cells (9 organs), endothelial cells (13 organs), lymphoid cells (7 organs), and myeloid cells (10 organs), rather than by batch or individually. For example, cell types related to development and function, such as diverse blood cells, secretory cells, PNS neurons, and CNS neurons, were also co-localized.
[0312] An important question in developmental biology is which transcription factors (TFs) are involved in producing this diverse array of cell types from an invariant genome. Next, we sought to leverage the breadth of this human cell atlas of chromatin accessibility to systematically evaluate TF motifs that are differentially accessible and thus nominate major regulators of cell fate in the context of human development in vivo.
[0313] As a first approach, we used a linear regression model to find the TF motifs found at the accessible sites of each cell that best explain the relationships between cell types. First, each tissue was treated independently, and in each of the 124 annotated cell type clusters, the most highly enriched motifs / TFs were identified from the JASPAR database, revealing both known regulators and potentially novel regulators. For example, in the placenta, the motif for SPI1 / PU1 (a well-established regulator of myeloid growth) was highly enriched at the peaks of myeloid cells, the motif for TWIST-1 (required for the formation of stromal progenitor cells) was enriched at the peaks of stromal cells, and the FOS::JUN motif was associated with chromatin accessibility in the extravillous trophoblast (a cell type in which the corresponding AP1 complex has been described as being specifically active).
[0314] Interestingly, unannotated clusters within the placenta are highly enriched for the GATA1::TAL1 motif (a well-established regulator of erythropoiesis). These cells cluster with erythroblasts from other tissues in the global UMAP, and upon further examination, major erythroid marker genes showed specific promoter accessibility. In the NNLS-derived workflow, this cluster was not annotated. This is because erythroblast clusters were not detected in the placenta in scRNA-seq studies, probably because the placenta is one of the few tissues with more ATAC than RNA cells. Thus, motif enrichment can assist in cell type annotation when the major regulators of the cell type are known.
[0315] We repeated this analysis for 54 major cell types observed across all tissues, i.e., after excluding cell types that appear in multiple tissues. As expected, the top motifs remained consistent with tissue-specific analyses and the literature, e.g., SPI1 / PU1 in myeloid cells, CRX in retinal pigment and photoreceptor cells, MEF2B in cardiomyocytes and skeletal muscle cells (31), and SRF in endocardial cells and smooth muscle cells. Most motifs are enriched in only one or two cell types, but neuronal TF motifs such as OLIG2, NEUROG1, and POU4F1 are enriched in multiple neuronal cell types. Another notable exception is HNF1B, which is conventionally associated with kidney and pancreas development, and its motif is enriched in 13 cell types spanning specific epithelial and secretory cells.
[0316] POU2F1 is an example of a TF that has not been previously associated with specific developmental branches. Rather, it is an exception within the POU family, being widely expressed and not thought to control a specific trajectory. In contrast, we found that its motif is enriched in several neuronal cell types, at least in human fetal development. Further supporting this, POU2F1 is specifically expressed in those same cell types.
[0317] In an extension of this observation, we next sought to more generally confirm whether TFs are differentially expressed in patterns that match the differential accessibility of their motifs, by leveraging the companion scRNA-seq atlas. For example, looking across all cell types annotated to the same tissue in both datasets, the expression of the hematopoietic pioneer factor SPI1 / PU1 strongly correlates with the enrichment of its motif at accessible sites. Intriguingly, this analysis also revealed many TFs with a negative correlation between expression and motif enrichment. On closer inspection, these TFs tended to be repressors. For example, GFI1B has been shown to act as an important repressor in erythroid and megakaryocytic development by recruiting histone deacetylases upon motif binding and inducing chromatin closure, for example, at the fetal hemoglobin locus. Consistent with this, we observed that its expression negatively correlates with the enrichment of its motif at accessible sites.
[0318] When the inventors classified TFs as "activators" or "repressors" based on GO terms, they found that TF expression and motif accessibility tend to be positively correlated with annotated activators and negatively correlated with annotated repressors, and that the mode of action of unclassified TFs can be predicted using the correlation between motif enrichment and expression. Most of the exceptions can be explained by the absence or conflict of GO terms, and when a literature search is conducted, they fit into the categories predicted by the correlation values. Therefore, this type of analysis can provide a systematic approach for classifying TFs as activators or repressors. For example, NFATc3 is generally described as an activator, but the inventors' analysis shows a mode of action of suppression in the development of T cells that are highly expressed but have a depletion of motifs at accessible sites. Such a mode of suppression of NFATc3's action has been suggested in the previous literature. Apart from the general classification, insights can also be obtained into the context of cell types in which TFs can act variably as activators or repressors. For example, it has been suggested that TFs such as FOXO3 act as activators in their unmodified state but act as repressors when phosphorylated, which could explain a more ambiguous relationship between expression and accessibility.
[0319] The above approach has the advantage of potentially associating known TFs systematically with new roles and not relying on pre-selecting differentially accessible sites for each cell type, and also the further advantage that the expression of a TF can be associated with the accessibility of its corresponding motif. However, it is limited in that it relies on a database of known TF motifs. As a different approach, a specificity score was also calculated for each accessible site, 2,000 of the most specific peaks were selected for each cell type, and de novo motif enrichment within this set was searched for by comparison with a CpG-matched background genomic sequence. Generally, the top new motifs for individual cell types match the top known motifs identified by linear regression. Interestingly, some cell types that did not have strong matches to known motifs (e.g., endothelial cells, stromal cells, Schwann cells) were still strongly associated with new motifs. In particular, for endothelial cells, such results will be further explained below.
[0320] Tissue cross-section analysis of blood cells and endothelial cells
[0321] The nature of this dataset creates an opportunity to investigate organ-specific differences in chromatin accessibility within widely occurring cell types, such as blood cells and endothelial cells. In the first pass of blood cell type annotation, myeloid cells, lymphocytes, erythroblasts, megakaryocytes, and hematopoietic stem cells could be distinguished. By extracting these blood lineages from all organs and reclustering, it became possible to additionally identify macrophages, B cells, NK / ILC3 cells, T cells, and dendritic cells, again employing an RNA-assisted annotation approach (notably, additional doublet cleaning steps are required to analyze similar cell types from multiple tissues; see "Methods"). Macrophages can be further classified into groups related to tissue origin, as well as phagocytic macrophages, as previously observed. This latter group was primarily identified in the spleen, followed by the liver and adrenal glands. Of particular interest in the blood lineages is the erythroblast, which is due to the spatio-temporal dynamics of erythropoiesis during fetal development. The inventors first detected this lineage in the liver, adrenal glands, heart, and placenta and further identified erythroblasts in the spleen, which was shallowly profiled by tissue cross-section analysis (initially annotated with only megakaryocytes and myeloid cells). The erythroblast ratio within the blood lineages of tissues was highest in the liver (consistent with this organ being the major site of erythropoiesis at this developmental stage), followed by the spleen and adrenal glands, mimicking the trend observed in the RNA data. The unexpected result that the adrenal gland is a potential site of fetal hematopoiesis will be further discussed in Example 2.
[0322] Further investigation of erythroblasts revealed that regions proximal to both the adult β-globin gene and the fetal γ-globin gene are accessible at this developmental stage, while the promoter of the embryonic ε-globin gene is inaccessible. The erythroblast cluster can further differentiate into five major Louvain clusters with differential chromatin accessibility, such as distinct erythroblast progenitor cell clusters. Accessible sites within the erythroblast progenitor cell cluster, as well as the adjacent early erythroblast cluster (erythroblast_3), are enriched for GATA1:TAL1 and other GATA motifs. By comparing the expression levels of various GATA factors in erythroblast progenitors, GATA1 / 2 can be nominated as a TF highly likely to be involved in the enrichment of this motif. Other erythroblast clusters corresponding to the late stages of erythropoiesis show motif enrichment for NFE2 / NFE2L2 (erythroblast_1) and KLF factors (erythroblast_2 / 4), and notably, the absence of enrichment of GATA motif accessibility is prominent. A recently published study on scRNA-seq regarding the mouse hematopoietic system reported that GATA2 is induced early in erythropoiesis and then decreases, while GATA1 is stably expressed. In contrast, studies on sorted bulk human in vitro cultured erythrocyte populations revealed a decrease in GATA1 expression from progenitors to differentiated erythroblasts (consistent with observations in human fetal tissues), as well as an increase in KLF1 levels and NFE-2 levels in late-stage erythroblasts. This result further indicates that there may exist epigenetically distinct subpopulations of differentiated erythroblasts whose accessibility landscapes are shaped by non-GATA factors such as KLF1 or NFE-2. For example, a distal regulatory element upstream of GYPA, which is used as an erythrocyte invasion receptor by Plasmodium falciparum, is most accessible in erythroblast_1 and contains a motif similar to the NFE-2 motif.
[0323] Another interesting tissue-crossing system is the vascular endothelium. Interestingly, there is no TF that is exclusively expressed in endothelial cells, suggesting that the endothelial-specific transcriptome is under combinatorial control by several TFs with overlapping expression in the endothelium. Consistent with this, no strong enrichment in endothelial cells was observed in the JASPAR motif analysis. On the other hand, the discovery of new motifs at the 2,000 most endothelial-specific peaks revealed strong enrichment across the background genomic sequences of motifs similar to ERG and SOX15. Since these motifs are not limited to endothelial cells (the ERG motif is more enriched in megakaryocytes and SOX15 is enriched in several cell types), and since the expression of these TFs is not limited to this cell type, they tended not to be weighted as strongly in our linear modeling approach. Thus, although ERG has already been described as a major regulator of endothelial function, it also promotes the conversion of culture to megakaryocytes.
[0324] Endothelial cells are present in all organs and need to perform both structural and highly differentiated functions, such as gas exchange in the lungs or fluid filtration in the kidneys. In this study, endothelial cells from 13 out of 15 organs were detected (the exceptions were the cerebellum and the eye, which were more shallowly profiled). When these cells were extracted across organs and reclustered, they were significantly separated according to tissue origin, in contrast to the erythroid lineage, despite a rigorous iterative filtering process (method) to remove any residual contaminating doublets. This also allowed us to observe tissue-specific programs of gene expression, as described in Example 2. Indeed, the peaks of most accessible chromatin closest to these differentially expressed genes have higher specificity scores in the tissues that match the ATAC data. Furthermore, endothelial cells from almost all organs showed enrichment of specific TF motifs. Notably, many of the TFs with enriched motifs are differentially expressed in the tissues that match the RNA data.
[0325] Collectively, these findings indicate that the general program of chromatin accessibility and gene expression in endothelial cells, a widely distributed cell type that must fulfill both general and organ-specific functions, is mediated by a combination of structural TFs such as ERG and SOX15, and tissue-specific TFs that promote further specialization. These analyses also highlight the merits of combining both de novo motif enrichment at specific peaks and a linear modeling approach across the tissue to nominate the key regulators underlying the chromatin accessibility landscape of individual cell types.
[0326] Another interesting example includes PAEP_MECOM-positive cell types in the placenta, which are identified in both the scRNA-seq atlas and the sc-ATAC-seq atlas. The regulatory regions of this lineage are strongly enriched for motifs of HNF1B, a factor previously associated with kidney and pancreas development. For example, HNF1B is expressed very specifically in the PAEP_MECOM cell lineage within the placenta. Due to the nature of ATAC-seq data that captures some genomic reads across the chromosome even at inaccessible sites, it is possible to distinguish the sex of cells based on reads from the Y chromosome or autosomes on the X chromosome. Interestingly, the inventors found that the PAEP_MECOM and IGFBP1_DKK-positive placental cell types, and to a lesser extent placental myeloid cells, have a significantly lower ratio of Y chromosome reads in male fetuses. Consistent with what is known about PAEP (glycodelin) and IGFBP1, these cell types may correspond to maternal endometrial epithelial and stromal cells, respectively.
[0327] CICERO
[0328] As a resource for further research, the inventors generated a Cicero core accessibility score and a Cicero gene activity score for each tissue of the dataset. The Cicero core accessibility score can be used to predict cis-regulatory interactions between accessible elements. The inventors combined elements paired by positive core accessibility scores to create a database of putative cis-regulatory interactions. This database contains 80 million unique core accessible pairs, including 4.5 million (6%) promoter-terminus pairs, 76 million (94%) terminus-terminus pairs, and 128,000 (0.2%) promoter-promoter pairs. The inventors found an average of 33 million core accessible pairs per tissue. 38% of the pairs were unique to a single tissue only, and only 0.007% of the pairs were detected in all 16 tissues. Pairs detected in more tissues were more likely to be promoter-terminus and promoter-promoter. The generated core accessibility scores and gene activity scores can be downloaded from the inventors' website.
[0329] Notably, 89% of the 436,206 sites initially identified were significantly differentially accessible (DA) at a 1% false discovery rate (FDR) in at least one of these 85 cell clusters compared to a control set of 2,040 cells (120 cells randomly sampled from each of 17 samples, see Additional Materials). To identify DA sites with accessibility restricted to specific clusters, a metric for quantifying gene expression specificity in scRNA-seq studies was adapted to chromatin accessibility and calculated for all 436,206 sites by all 85 clusters. 39% (167,981 / 436,206) of the accessible sites were classified as cluster-restricted (i.e., increased accessibility in a limited number of clusters), and 55% (92,334 / 167,981) of these were restricted to a single cluster.
[0330] Implications of cell types in common human traits and diseases
[0331] Most of the heritability of common human traits and diseases, as measured by genome-wide association studies (GWAS), is partitioned into distal regulatory elements, which are often cell type specific. As a result, most studies are spent intersecting GWAS signals with bulk DNase hypersensitivity data (and other epigenetic features) for the purpose of systematically associating a particular disease with the dysfunction of a particular tissue. However, the resolution of such studies is severely limited by cell type heterogeneity. The inventors considered whether, given the conservation of chromatin accessibility between mouse and human, the data could be used to gain further understanding of the cell type-specific effects of the various genes underlying complex human traits, regardless of species differences. Therefore, despite the fact that the inventors' data were generated in mouse tissues, they sought to apply state-of-the-art methods to detect cell type-specific enrichment of human heritability.
[0332] To do this, we used partitioned linkage disequilibrium (LD) score regression (LDSC) to quantify the enrichment of heritability of human traits within DA peaks for each of the 85 clusters. After transferring human SNPs to orthologous coordinates in the mouse genome, we calculated the enrichment of heritability for 32 phenotypes across the DA peaks obtained for each of the 85 clusters. Of the 85 cell types, 55 had enrichment for at least one phenotype, and of the 32 phenotypes, 28 were enriched for at least one cell type. As a major trend, strong enrichment of heritability for autoimmune diseases such as lupus, celiac disease, and Crohn's disease was observed within clusters corresponding to leukocytes, while enrichment for neurological traits such as bipolar disorder, educational attainment, and schizophrenia occurred in neuron cell types. Notably, most of these enrichments were not prominent at peaks from bulk tissues, demonstrating cell type values defined by single-cell chromatin accessibility data. Many of the enrichments were as expected. For example, the strongest enrichment of heritability for low-density lipoprotein (LDL) cholesterol, high-density lipoprotein (HDL) cholesterol, and triglycerides was present in hepatocytes, but interestingly, LDL cholesterol was also significant in the Henle's loop kidney epithelium. Similarly, the strongest enrichment of heritability for immunoglobulin A (IgA) deficiency was present within a cluster of T cells. These signals can also provide further understanding of the importance of cell subtypes. As an example of this trend, enrichment of heritability for bipolar disorder has been observed for multiple neuron clusters, but the strongest enrichment is with excitatory neurons. In contrast, the heritability for Alzheimer's disease is not enriched in any class of neurons. Instead, its strongest enrichment is found in a cluster of microglia.
[0333] To extend the inventors' analysis to larger trait sets, summary statistics from genome-wide association studies (GWAS) for 2,419 traits of over 300,000 individuals were downloaded from the UK Biobank (nealelab.github.io / UKBB_ldsc / ). Focusing on 405 traits with effective sample size ≥5,000 and estimated heritability ≥0.01, significant enrichment of heritability was observed in 273 traits in at least one cell type, and 74 out of 85 cell types showed enriched heritability for at least one trait. For autoimmune and neurological traits, the same large trends as described above are seen here, but the far larger number of traits measured by the UK Biobank reveals further trends. For example, numerous measurements of body size and composition (e.g., body mass index) are also related to cell types in the brain (Figure 18B). In addition, specific subsets of T cells (12.1, 12.2) are more strongly associated with asthma and allergic rhinitis than other cell types such as other T cell clusters. At a more refined level, heart attacks are related to endothelial cells from the liver (25.3), while endothelial cells from other endothelial clusters are not. On the other hand, gout is associated with kidney proximal tubule cells. The framework demonstrated herein can be readily applied to single-cell chromatin accessibility data collected from any human or mouse tissue and any heritable trait.
[0334] One result of the new design is its compatibility with both two-level ("2lv2", i.e., "two-level version 2 protocol") and three-level ("3lv2") configurations, providing additional flexibility in the test design (Figure 9).
[0335] Finally, various conditions for fixing cells or nuclei with formaldehyde were tested to enable long-term stable storage. The inventors found that choosing the buffer used for fixation and the isolation of nuclei before or after fixation presents a choice between complexity and specificity. In the current study, the inventors selected a fixation protocol that increases complexity / sensitivity at the expense of specificity, which can be determined by the end-user of the protocol.
[0336] Materials and Methods
[0337] Cell Culture
[0338] Gm12878 cells were cultured and maintained in RPMI 1640 medium (Thermo Fisher catalog number 11875-093) containing 15% FBS (Thermo Fisher catalog number SH30071.03) and 1% Pen-strep (Thermo Fisher catalog number 15140122). These were counted and split three times a week at 300,000 cells / mL. The CH12-LX mouse cell line was provided by the Michael Snyder lab (Stanford). The cells were cultured in RPMI1640 medium containing 10% FBS, 1% Pen-strep (penicillin and streptomycin), and 1×10^5 M B-ME. These were counted and maintained at a density of 1×10^5 cells / mL and split three times a week to maintain the cell concentration. Both cell lines were incubated at 5% CO2 and 37°C.
[0339] Nuclei Isolation and Fixation from Cell Lines
[0340] For suspension cells, ~10 - 100 million cells are obtained and the cells are pelleted by spinning at 500 x g for 5 minutes at room temperature. The supernatant is aspirated and 1 mL of Omni-ATAC lysis buffer (10 mM NaCl, 3 mM MgCl2, 10 mM Tris-HCl Resuspend the pellet in pH 7.4, 0.1% NP40, 0.1% Tween 20 and 0.01% digitonin), and incubate on ice for 3 minutes. Add 0.1% Tween 20 to 5 mL of 10 mM NaCl, 3 mM MgCl2, 10 mM Tris-HCl pH 7.4, and pellet at 500 x g, 4 °C for 5 minutes. Aspirate the supernatant and resuspend the nuclei in 5 mL of 1X DPBS (Thermo Fisher catalog number 14190144). To crosslink the nuclei, add 140 μL of 37% formaldehyde to methanol (VWR catalog number MK501602) in one addition, and the final concentration was 1%. Incubate the fixation mixture at room temperature for 10 minutes, inverting every 1 - 2 minutes. To quench the crosslinking reaction, add 250 μL of 2.5 M glycine, incubate at room temperature for 5 minutes, then incubate on ice for 15 minutes to completely stop the crosslinking. Place 20 μL of the quenched crosslinking mixture into 20 μL of trypan blue for counting. Spin the crosslinked nuclei at 500 x g, 4 °C for 5 minutes and aspirate the supernatant. Resuspend the fixed nuclei in an appropriate amount of lysis buffer (50 mM Tris, pH 8.0, 25% glycerol, 5 mM Mg(OAc)2, 0.1 mM EDTA, 5 mM DTT (Sigma-Aldrich catalog number 646563-10X0.5 mL), 1× protease inhibitor cocktail (Sigma-Aldrich catalog number P8340) to obtain 2 million nuclei per 1 mL aliquot, snap freeze in liquid nitrogen and store at -80 °C.
[0341] Procurement and storage of tissues
[0342] Isolate the target tissue, wash it away with 1X HBSS (containing Ca. and Mg.), and then absorb and dry it on semi-wet gauze. Place the dried tissue on sturdy foil or in a cryotube, and rapidly freeze the tissue using liquid nitrogen. Store the frozen tissue at -80 °C.
[0343] Nuclei isolation and fixation of frozen fetal tissues
[0344] On the day of grinding, place a cloth towel between the dry ice and the metal, and pre-cool the pre-labeled tube and hammer on the dry ice. Use a 18-inch by 18-inch sturdy foil to make a "stuffing", fold it in half twice to make a rectangle. Then fold it in half twice more to make a square. Put the frozen tissue inside the foil "stuffing", and then place the tissue in the foil stuffing inside a pre-cooled 4mm plastic bag to prevent the tissue from falling onto the dry ice if the foil ruptures. Cool this tissue packet between two pieces of dry ice. Use a pre-cooled hammer to manually grind the tissue inside the packet. Avoid the grinding operation with 3 - 5 impacts and take breaks to prevent the sample from heating up. Cool the hammer as needed until the tissue is uniform and repeat the grinding. Divide the ground tissue equally into pre-labeled and pre-cooled 1.5 mL LoBind and nuclease-free snap-cap 1.5 mL tubes (Eppendorf catalog number 022431021). The aliquots of the powdered tissue can be stored at -80 °C until further processing.
[0345] On the day of nuclear isolation, add lysis buffer directly to the tube or place the frozen aliquot into a 60 mm dish containing cell lysis buffer and further mince with a blade. The aliquot of powdered tissue should be easily retrievable from the storage tube without sample loss, as long as the aliquot does not thaw at any point during storage. The inventors estimated ~20,000 cells per 1 mg of original tissue weight, and performance may vary from tissue to tissue. Resuspend the minced tissue in 1 mL of Omni lysis (RSB + 0.1% Tween + 0.1% NP-40 and 0.01% digitonin), then transfer to a 15 mL Falcon tube. Incubate the nuclei on ice for 3 minutes, then add 5 mL of RSB + 0.1% Tween20. Centrifuge the nuclei at 500 x g for 5 minutes at 4°C. Aspirate the supernatant and resuspend in 5 mL of 1X DPBS. Pass the nuclei in 1X DPBS through a 100 micron cell strainer (VWR catalog number 10199-658) to remove tissue chunks. In the draft, crosslink the nuclei by adding 140 uL of 37% formaldehyde to the methanol in one addition to a final concentration of 1% and quickly mix by inverting the tube several times. Incubate at room temperature for exactly 10 minutes, gently inverting the tube every 1-2 minutes. Add 250 uL of 2.5 M glycine (freshly made and filter sterilized) to quench the crosslinking reaction and mix well by inverting the tube several times. Incubate at room temperature for 5 minutes, then on ice for 15 minutes to completely stop the crosslinking. Use a hemocytometer to count the nuclei and confirm the final amount of freezing buffer to add. The goal is to freeze ~1-2 million nuclei / tube. Centrifuge the crosslinked nuclei at 500 x g for 5 minutes at 4°C, aspirate the supernatant, and resuspend the pellet in 1-10 mL of freezing buffer supplemented with 1x protease inhibitor and 5 mM DTT. Rapidly freeze the nuclei in liquid nitrogen and store the nuclei at -80°C.
[0346] Processing of sciATAC-seq3 samples (library construction and qc)
[0347] Remove the frozen fixed nuclei from -80°C and place them on a dry ice bed. Thaw the nuclei in a 37°C water bath until thawed (~30 seconds to 1 minute), and transfer the nuclei to a 15 mL Falcon tube. Pellet the nuclei at 500 x g for 5 minutes at 4°C. Aspirate the supernatant without disturbing the pellet, resuspend the pellet in 200 μL of Omni lysis buffer, and then incubate on ice for 3 minutes. Wash the lysis buffer with 1 mL of ATAC-RSB containing 0.1% Tween20, and gently invert the tube 3 times to mix. Take 20 μL of nuclei and 20 μL of trypan blue to count the nuclei. While counting, maintain the nuclei on ice as much as possible in the future. For 3-level indexing experiments in 384^3d, the number of nuclei input is 4.8 million per well per tissue @50,000 nuclei, or samples spread over 96 reactions. Pellet the nuclei and resuspend them in the pre-made tagging reaction master mix (Nextera TD buffer, 1X DPBS, 0.1% digitonin, 0.1% Tween 20, and water). Using a wide-bore tip (Rainin Instrument Co catalog number 30389249) across the entire LoBind 96-well plate (Eppendorf catalog number 30129512), aliquot 47.5 μL of nuclei in the tagging mix. Add 2.5 μL of Nextera v2 enzyme (Illumina Inc catalog number FC-121-1031), seal the plate with adhesive tape, and rotate at 500 x g for 30 seconds. Incubate the plate at 55°C for 30 minutes to tag the DNA. Add 50 μL of stop reaction mixture (40 mM EDTA containing 1 mM spermidine) to stop the tagging reaction, and then incubate at 37°C for 15 minutes. Using a wide-bore tip, pool the tagged nuclei, pellet them at 500 x g for 5 minutes at 4°C, and then wash with ATAC-RSB containing 0.1% Tween20. Pellet the nuclei at 500 x g for 5 minutes at 4°C, aspirate the supernatant, and resuspend in 384 μL of ATAC-RSB containing 0.1% Tween 20. PNK reaction master mix (1X Prepare PNK buffer (NEB catalog number M0201L), 1 mM rATP (NEB catalog number P0756S), water, and T4 polynucleotide kinase (NEB catalog number M0201L), and add them to the nuclei. Aliquot 5 μL of the PNK reaction mix into four LoBind 96-well plates, seal with adhesive tape, and rotate at 500 x g at 4 °C for 5 minutes. Incubate the PNK reaction at 37 °C for 30 minutes. Add 13.8 μL of ligation master mix (1X T7 ligase buffer (NEB, catalog number M0318L), 9 μM N5_sprint (IDT), water, and 2.5 μL of T7 DNA ligase enzyme (NEB catalog number M0318L) directly to the PNK reaction. Use a multi-channel, i.e., 96-head dispenser (Liquidator, catalog number 17010335) to dispense 1.2 μL of 50 μM into each well across four 96-well plates Add N5_oligo (IDT). Seal with adhesive tape, rotate at 500 x g for 30 seconds, then incubate at 25 °C for 1 hour. After the first ligation, add 20 μL of 40 mM EDTA containing 1 mM spermidine to stop the ligation reaction and incubate at 37 °C for 15 minutes. Using a wide-mouth pipette tip, pool each well into a trough and transfer to a 50 mL Falcon tube. Pellet the nuclei at 500 x g at 4 °C for 5 minutes, aspirate the supernatant, and resuspend the nuclei in 1 mL of ATAC-RSB containing 0.1% Tween 20 to wash away all remaining ligation reaction mix. Pellet the nuclei at 500 x g at 4 °C for 5 minutes and aspirate the supernatant without disturbing the pellet. Prepare the N7 ligation master mix (1X T7 ligase buffer, 9 μM N7_sprinter (IDT), water, and T7 DNA ligase) and resuspend the nuclei in the ligation master mix. Transfer the nuclei suspended in the master mix to a trough and, using a wide-mouth pipette tip, aliquot 18.8 μL of the ligation master mix into four 96-well LoBind plates, then add 1.2 μL of 50 μM N7_oligo (IDT) to each well across the four 96-well plates. Seal the plates with adhesive tape, rotate at 500 x g for 30 seconds, then incubate at 25 °C for 1 hour, then add 20 μL of 40 mM EDTA and 1 mM spermidine to stop the ligation and incubate at 37 °C for 15 minutes. Pool the wells into a trough using a wide-mouth pipette tip and then transfer to a 50 mL Falcon tube. Pellet the nuclei at 500 x g at 4 °C for 5 minutes, aspirate the supernatant, and resuspend the nuclei in 2 mL of Qiagen EB buffer (Qiagen catalog number 19086). Obtain 20 μL of resuspended nuclei and 20 μL of trypan blue and count the nuclei. Dilute the nuclei to 100 - 300 nuclei / μL and aliquot 10 μL / well into four 96-well LoBind plates. To reverse crosslink the nuclei, prepare a reverse crosslink master mix of EB buffer, proteinase k (Qiagen, catalog number 19133), and 1% SDS (1 μL / 0.5 μL / 0.5 μL / well respectively) and add 2 μL to the nuclei in each well.Seal with tape, spin at 500 x g for 30 seconds, and incubate at 65 °C for 16 hours. Perform test PCR amplification and monitor the reaction with SYBR Green in several wells of the plate to determine the optimal number of cycles. Based on the test PCR results, amplify the remainder of the reverse crosslinked plate with 7.5 uL of NPM, 0.5 uL of BSA (NEB, catalog number B9000S), 1.25 uL of indexed P5_10 uM (IDT), 1.25 of indexed P7_10 uM (IDT), and water per well. Depending on the tissue and nuclei recovery after two ligations, 11 - 13 cycles are typical for the inventors. The cycle conditions were 3 minutes at 72 °C, 30 seconds at 98 °C, "10 seconds at 98 °C, 30 seconds at 63 °C, 1 minute at 72 °C" for 11 - 13 cycles, and hold at 10 °C. Pool the amplification products from the 96 well plate into a trough, purify using Zymo Clean&Concentrate-5 (Zymo Research catalog number D4014) according to the manufacturer's specifications, and divide into 4 columns. Elute each column in 25 uL of EB buffer and then combine into one tube. Add 100 uL of AMPure beads (Agencourt, catalog number A63882) to the purified PCR product to further remove any remaining primer dimers and follow the manufacturer's purification process. Elute the final library from the beads in 25 uL of Qiagen EB buffer. Quantify the final library using the D5000 ScreenTape (Agilent catalog number 5067-5588 ScreenTape, 5067-5589 reagent) and the Agilent 4200 Tapestation System which establishes a 200 - 1000 base pair window and measures the nM concentration of the fragments that cluster the wells during sequencing. Make a 2 nM pool from equimolar pooling and sequence at a loading concentration of 1.8 pM using the custom recipe and primers of the NextSeq High Output 150 cycle kit (Illumina catalog number 20024904).
[0348] Data processing for method development
[0349] The data processing of the chicken experiment conducted to develop sci-ATAC-seq3 was performed as described above. Briefly, BCL files were converted to fastq files using bcl2fastq v2.16 (Illumina). Each read was associated with a cell barcode consisting of four components. At the P5 end of the molecule, there were row addresses added for tagging and PCR, and at the P7 end of the molecule, there were column addresses added for tagging and PCR. To correct the errors of these barcodes, the inventors split them into these four components and corrected them to the closest barcode within an edit distance of 2 as long as the correction was unique at the required edit distance. If any of the four barcodes could not be corrected to a known barcode, the corresponding read pair was dropped. Subsequently, the reads were adjusted using Trimmomatic with the option "ILLUMINACLIP:{adapters_path}:2:30:10:1:true TRAILING:3 SLIDINGWINDOW:4:10 MINLEN:20". Then, the adjusted reads were mapped to the hybrid human / mouse (hg19 / mm9) genome using bowtie2 with the option "-X 2000 - 3 1". Subsequently, reads that were not mapped to appropriate pairs in the genome with at least 10 accuracy were filtered out and removed using samtools with the option "-f3 - F12 - q10", and only the reads mapped to autosomes or sex chromosomes were retained for downstream analysis. Duplicate reads were removed for each cell barcode using a custom script. Note that read pairs are not maintained as duplicates, unlike the tissue pipeline (discussed below).
[0350] Data processing for tissue samples
[0351] The method for processing sequencing data from tissue samples follows faithfully the methods used and has many optimizations for scaling to larger datasets, but for the sake of expediency, is described herein. The BCL files were converted to fastq files using bcl2fastq v2.20 (Illumina). Reads with modified barcodes included in the read names were written to separate R1 / R2 files for each sample within our dataset. All mismatches mapping to a known barcode set were precomputed (feasible because barcode lengths are short and relatively few), and a correction script was run using pypy, an alternative to the cpython interpreter which is extremely fast for this particular task, and this calculation was parallelized across different lanes of the sequencing run. This comprehensively improved the runtime over previous methods significantly.
[0352] Next, Trimmomatic was used with the option "ILLUMINACLIP:{adapters_path}TRAILING:3 SLIDINGWINDOW:4:10 MINLEN:20" to trim low-quality base / adapter sequences from the 3' end, then bowtie2 was used with the option "-X 2000 3 1" to map the trimmed reads to the hg19 reference genome, and then read pairs that did not map uniquely to autosomes or sex chromosomes with at least 10 mapping quality were filtered out and removed using Samtools --samtools view -L{whitelist of chromosomes} -f3 -F12 -q10 -bS. The resulting BAM files were sorted, the aligned reads for each sample were merged using sambabamba, and the resulting BAM files were indexed. This process was parallelized across samples / lanes as much as possible, but by providing more threads for trimmomatic / bowtie2 / sambabamba for each process, the runtime will be improved.
[0353] Subsequently, PCR duplicates within cells were identified by identifying a unique set of fragment endpoints within each cell. In our previous studies, the resulting duplicate BAM files did not always maintain the correct read names between read pairs written to the duplicate BAM files (for each unique fragment, representative reads for R1 and R2 were randomly selected independently), which was the cause of compatibility issues with some tools such as SnapATAC (github.com / r3fang / SnapATAC). We corrected this issue and also wrote files that exactly mirror 1) the BED file of fragment endpoints per cell, and 2) the fragments.tsv.gz file provided by 10x Genomics for the scATAC solution.
[0354] For each sample, the BED file of the unique fragment endpoints per cell was used for peak calling of each sample by MACS2--macs2 callpeak-t{bed}-f BED-g hs--nomodel--shift-100--extsize 200--keep-dup all--call-summits-n{sample_name}-o{output_dir}. The resulting {outdir} / {sample_name}_peaks.narrowPeak files were sorted and output as BED files. Peak calls from all samples included in the downstream analysis (additionally excluding our standard) were merged using bedtool to form a master set of peaks. As previously explained, it was intentional to use BED files for peak calling herein, and it was noted that the behavior of macs2 for BAM input was not considered. When using a BAM file as input, MACS2 either discards one of the read pairs that uses R1 / R2 independently (effectively downsampling the input data), or uses the entire insert during coverage calculation if the BAM file is explicitly specified as a paired-end (we desire to calculate coverage only at the endpoints, not along the entire insert). By using a BED file, all data can be used and coverage can be calculated using only the windows around the molecular endpoints.
[0355] Furthermore, for each sample, 1) a sparse matrix was created that counts the reads that enter the master set of peaks, and 2) the reads that enter the gene body extended by 2 kb upstream and the 5 kb genomic window. Additionally, the total number of reads per cell from the annotated TSSs (+ / - 1 kb around each TSS), the ENCODE blacklist regions, and the merged peak sets for QC purposes were listed.
[0356] Also, peaks were constructed by motif matrices using the method used in the 10x Genomics scATAC pipeline (see support.10xgenomics.com / single-cell-atac / software / pipelines / latest / algorithms / overview). Briefly, the method from 10x calculates the GC% distribution of peaks and bin peaks into equal percentile ranges of GC content, allowing for the separate discovery of motif occurrences within each bin. The MOODS package was used to identify background nucleotide compositions matched to each GC bin for motif occurrences and to mitigate GC bias for motifs in the JASPAR motif database at a p-value threshold of 1E-7. These hits are used to build motifs by a peak matrix that can be used to calculate the motif matrix by cell count in downstream analysis. This matrix is binarized such that only one instance of a motif can be counted per peak.
[0357] Cell barcodes were separated from the background barcode distribution using a modified version of the method used in the 10x Genomics scATAC pipeline (see link above). Briefly, it was fit to two negative binomial (noise vs. signal) mixtures. Instead of the method used by 10x, the k-means method was applied to the log-scaled total fragment count distribution to establish an initial threshold between these two distributions, and the maximum value of the cluster with the lower average total count was taken as the initial threshold. This initial threshold was used to determine the starting parameters of the two distributions using the maximum likelihood estimate and further refined by the expectation maximization approach. As described in 10x, this fit can be improved by applying a left shift to the count distribution. Different from the 10x method, this shift was determined by trying several shifts from 2 to 12 to obtain a mixture distribution model with the best fit. Finally, in contrast to the 10x approach, this method was applied to the total fragment count distribution rather than the count distribution within the called peaks. The final threshold selected was the minimum number that resulted in an odds ratio of 20 or more (beneficial to the signal) and removed at least 0.5% of the signal distribution as estimated from the CDF of the signal distribution (the inventors found that this second criterion prevented fitting to thresholds that would otherwise appear overly ambiguous).
[0358] Cell-level QC, dimensionality reduction, and clustering
[0359] As described above, the total number of unique reads falling within the region surrounding the TSS (+ / - 1 kb) in the peaks and ENCODE blacklist regions was tabulated for each cell. Using these totals, for each sample, sample-specific cutoffs for the percentage of unique reads in the peaks and the percentage of unique reads falling within the TSS, as well as a global cutoff of 0.5% of the unique reads obtained from the ENCODE blacklist regions, were selected by visual inspection of these distributions. Due to a few samples having an automatically determined threshold significantly lower than the other samples in the dataset, a global threshold of 1000 unique reads per cell (or 500 unique fragments per cell) was applied to raise the automatically determined threshold for the corresponding samples. The previously developed nucleosome binding scores were examined, but as previously observed for mouse testis, no distinct distribution of outliers was observed, so these scores were not used in the QC. Prior to downstream processing, peaks overlapping the ENCODE blacklist regions or corresponding to the sex chromosomes were removed (the latter to avoid introducing potential batch effects between samples of different sexes). Also, peaks exceeding two standard deviations from the mean of the log-scale counts per peak distribution were excluded to remove peaks with very low counts within the tissue being analyzed.
[0360] All downstream processing was performed on one tissue at a time by pooling the cells passing through all samples of a given tissue.
[0361] After filtering, a modified version of the Scrublet algorithm was used to remove cells that were most likely to be doublets. Briefly, peaks by the cell matrix were used to simulate doublets as the sum of randomly selected cells from the dataset. Next, LSI was performed as described below using the original cell matrix and the simulated doublets. Note that in this step, similar to how Scrublet applies the magnification from the original dataset of scRNA-seq data, inverse document frequency (IDF) terms obtained from the original dataset without using the simulated doublets were used. The nearest neighbors of each cell were found in the resulting 50-dimensional space, and the proportion of pseudo-doublets in the neighborhood was calculated as the doublet score. The top 10% of the cells within each sample with the highest doublet scores were excluded.
[0362] Regarding dimensionality reduction, first, it was found that the latent semantic indexing (LSI; alternatively, latent semantic analysis, i.e., LSA) described so far does not function well with the data collected in this study. It was judged that this might be due to sparsity, and several alternative methods such as CisTopic and SnapATAC were investigated. Each of these methods initially seemed to function better than LSI. Initially, even considering the fundamental similarities of these methods and the nature of the data, the reason for such a state was unclear. The inventors discovered that a simple logarithmic scaling of term frequencies in LSI, which had not been done before, resulted in performance very similar to that of the other tools tested. This might be due to the exponential distribution of the total count per cell and the influence of strong outliers on the PCA step of LSI without logarithmic scaling. This is elaborated in andrewjohnhill.com / blog / 2019 / 05 / 06 / dimensionality-reduction-for-scatac-data / . Note that the observed differences in the use of logarithmic scaling are particularly large in sparse datasets with a large range of total counts per cell. Also, note that other groups have favorably compared LSI with all other existing methods for reducing the dimensions of scATAC to confirm the inventors' independent findings. Also, since very similar performance was observed when using genomic peaks or 5kb windows, the choice was made to use peaks as was mainly done in previous studies.
[0363] In summary, at a certain point in time, for each tissue, the LSI was performed with a binarized window on the cell matrix from all passing cells of that tissue, one tissue at a time. First, all positions of individual cells were weighted by the logarithm (total number of accessible peaks in the cell) (logarithmically scaled "term frequency"). Then, these load values were multiplied by the logarithm (1 + inverse frequency of each position across all cells), i.e., the "inverse document frequency". Then, singular value decomposition was used on the TF-IDF matrix to generate a lower-dimensional representation (PCA) of the data, retaining only the 2nd to 50th dimensions (since the 1st dimension tends to be highly correlated with the lead depth). Then, L2 normalization was performed on the PCA matrix to further account for the difference in the number of unique fragments per cell. This L2-normalized PCA matrix was used in all downstream steps.
[0364] No evidence of significant batch effects between samples was observed, but the Harmonious batch correction algorithm was applied to the PCA space to correct for batch effects between different samples. Harmony was selected mainly because it can be easily scaled to large datasets and can use existing PCA coordinates.
[0365] This corrected L2-normalized PCA space was used as input to Louvain clustering and UMAP, as performed in Seurat V3.
[0366] Specificity score
[0367] Before calculating the specificity score, all peaks overlapping with the ENCODE blacklist regions were filtered out and removed. The specificity score was calculated for each site / cell type pair as described above.
[0368] Motif enrichment
[0369] Before calculating motif enrichment, all peaks overlapping with ENCODE blacklist regions were filtered out. First, a motif × cell matrix is obtained by multiplying the peak × motif matrix by the corresponding peak × cell matrix (summed over all cells within the subset of the target data as described above). Note that the dataset is downsampled to include a maximum of 800 cells per annotation (e.g., cell type) to reduce computational cost and reduce the overrepresentation of a very large number of cell types when calculating enrichment in downstream steps. Then, for each annotation, negative binomial regression is performed using the speedglm package, with two input variables, namely the annotation indicator column as the main variable of interest and the logarithm of the total number of non-zero entries in the input peak matrix per cell as the covariate, to predict the total motif count. The fold change of the motif count of the annotation of interest relative to cells from all other annotations, i.e., exp(intercept+annotation_efficient) / exp(intercept), is estimated using the coefficients and intercept of the annotation indicator column. This test is performed for all motifs in the entire population, and then the p-values are corrected using the Benjamini Hochberg procedure.
[0370] Example 2
[0371] Human Cell Atlas of Gene Expression in Development
[0372] Summary
[0373] The emergence and differentiation of cell types during human development are fundamentally fascinating. An assay for single-cell profiling of gene expression based on three levels of combinatorial indexing (sci-RNA-seq3) was applied to 121 fetal tissues representing 15 organs, profiling transcription in a total of 4 to 5 million single cells. From these data, cell types were identified and annotated with respect to marker genes, expression, and regulatory modules. In the initial analysis of these data, we focused on cell types spanning multiple organ systems, such as epithelial, endothelial, and blood cells. Interesting observations include organ-specific endothelial specialization, potential novel sites of fetal erythropoiesis, and potential novel cell types. Combined with an accompanying human cell atlas of chromatin accessibility during development, these data are a rich resource for exploring human biology.
[0374] The text
[0375] For several reasons, we set out to generate a human cell atlas of both gene expression and chromatin accessibility using tissues obtained during development. First, genetic diseases, which largely involve developmental components, account for a disproportionately high rate of pediatric morbidity and mortality. These include thousands of Mendelian disorders, in which both genetic and non-genetic factors contribute significantly, as well as more common diseases (such as congenital heart failure, other birth defects, neurodevelopmental disorders, etc.). A reference cell atlas generated from tissue development can serve as the basis for a tissue-based effort to understand the specific molecular and cellular events that underlie each of these pediatric diseases.
[0376] Second, developing tissues provide a much better opportunity to study the in vivo emergence and differentiation of human cell types than adult tissues. Compared to embryonic and fetal tissues, adult tissues are populated by differentiated cells and do not simply represent many cell states. With better resolution of the in vivo developmental trajectories, single-cell atlases generated from developing tissues can inform a fundamental understanding of in vivo human biology and, importantly, our understanding of cell reprogramming and cell therapy.
[0377] Third, for many adult human organs, pioneering cell atlases have already been reported, but the independent nature of these studies makes it difficult to investigate differences between cell types that occur in different tissues, such as epithelial cells, endothelial cells, and blood cells. Specifically, comparisons based on existing data are confounded by differences in sample processing and technical platforms between groups that generate organ-specific cell atlases.
[0378] Towards a human cell atlas of gene expression, we applied a recently developed assay for single-cell RNA-seq based on three levels of combinatorial indexing (sci-RNA-seq3) to 121 fetal tissues representing 15 organs and profiled gene expression in over 5 million cells in total (Figure 11). In Example 1, we describe the profiling of chromatin accessibility in 1.6 million cells from the same organs based on an overlapping set of samples. The organs profiled span diverse systems and are most notably absent for bone marrow, bone, gonads, and skin.
[0379] Samples were obtained from 28 fetuses in the estimated gestational age range of 72 - 129 days. Briefly, these were snap frozen, pulverized, and the resulting powder was split for different assays. For sci-RNA-seq3, nuclei were extracted directly from the cryogenic, lysed powder and then fixed with paraformaldehyde. In the kidney and digestive organs, which are rich in RNases and proteases, paraformaldehyde-fixed cells rather than nuclei were used to increase cell and mRNA recovery. For each experiment, nuclei or cells from a given tissue were deposited into different wells, thereby also identifying the source for the first index of the sci-RNA-seq3 protocol. As batch controls for experiments with nuclei, a mixture of human HEK293T and mouse NIH / 3T3 nuclei, or nuclei from a common "sentinel" tissue (also used in sci-ATAC-seq3 experiments) were placed into one or more wells. As batch controls for experiments with cells, cells from common pancreatic tissue (whose nuclei were also profiled) were placed into one or more wells.
[0380] Sci-RNA-seq3 libraries from 7 experiments over 7 runs of Illumina NovaSeq were sequenced, generating 68.6 billion reads in total. The data were processed as described above to recover 4,979,593 single-cell gene expression profiles (UMI > 250). Single-cell transcriptomes from human-mouse control wells were overwhelmingly species coherent (~5% collisions). Uniform Manifold Approximation and Projection (UMAP) of nuclei or cells from sentinel tissue showed that cell type differences overwhelm batch effects between any experiments. Integrated analysis of nuclei and cells corresponding to common pancreatic tissue using Seurat also yielded highly overlapping assignments.
[0381] The inventors profiled the median of 72,241 cells or nuclei per organ (maximum 2,005,512 (cerebrum), minimum 12,611 (thymus)). Compared to other large-scale single-cell RNA-seq atlases, comparable numbers of UMIs per cell or nucleus were recovered (median 863 UMIs and 525 genes), despite relatively shallow sequencing (~14,000 raw reads per cell). As expected, nuclei showed a higher proportion of UMI mapping to introns than cells (56% for nuclei, 45% for cells, p < 2.2e-16, two-sided Wilcoxon rank-sum test). Unless otherwise specified, "cells" is used to refer to both cells and nuclei.
[0382] Tissues were readily identified as being from male (n = 14) or female (n = 14) based on the expression of sex-specific genes. Each of the 15 organs was represented by at least two of each sex and multiple samples (median 8), such as a range of gestational ages. UMAP visualization of the "pseudo-bulk" transcriptome of each tissue clustered by organ, not by individual or experiment. Approximately half of the expressed protein-coding transcripts were differentially expressed across this set of pseudo-bulk transcriptomes (11,766 out of 20,033, FDR 5%).
[0383] Applying Scrublet, 6.4% putative doublet cells were detected, corresponding to a 12.6% doublet estimate that includes both intra- and inter-cluster doublets. The strategy previously developed for the 2 million mouse organogenesis cell atlas (MOCA) was then applied to remove low-quality cells, doublet-enriched clusters, and spike-in HEK293T and NIH / 3T3 cells. All analyses described below are based on 4,062,980 human single-cell gene expression profiles from 112 fetal tissues remaining after this filtering step.
[0384] Identification of 77 major cell types
[0385] After filtering for low-quality cells and doublet-enriched clusters, 4 million single-cell gene expression profiles were subjected to UMAP visualization and Louvain clustering by Monocle 3 on an organ basis. Overall, 172 cell types were first identified and annotated based on cell-type specific markers from the literature. Excluding annotations common to tissues, this decreased to 77 major cell types, of which 54 were observed in only a single organ (e.g., Purkinje neurons of the cerebellum), and 23 were observed in multiple organs (e.g., vascular endothelial cells of each organ). These 77 major cell types included a median of 4,829 cells and ranged from 68 cells (SLC26A4_PAEP-positive cells of the adrenal gland) to 1,258,818 cells (excitatory neurons in the brain). Each major cell type contributed to multiple individuals (median 9). Despite differences in species, developmental stage, and technology, we recovered almost all major cell types identified by previous atlas-building efforts targeting the same organs. We identified the median number of 12 major cell types per organ, which ranged from 5 (thymus) to 16 (eye, heart, and stomach). No correlation was observed between the number of profiled cells and the number of identified cell types (ρ = -0.10, p = 0.74).
[0386] On average, 11 marker genes per major cell type were identified (minimum 0, maximum 294; defining differentially expressed genes as those with at least a 5-fold difference in expression between the top and second-ranked cell types for expression; FDR 5%). Due to similar cell types in other organs (e.g., ENS glia and Schwann cells), there were some cell types without marker genes at this threshold. Therefore, we also report a set of "tissue marker genes" determined per organ using the same procedure (average 147 markers per cell type; minimum 12, maximum 778).
[0387] Canonical markers were generally observed and were actually important in this annotation process, but as far as is known, most of the markers observed are novel. For example, OLR1, SIGLEC10, and non-coding RNA RP11-480C22.1 are among the strongest markers of microglia, along with more established microglial markers such as CLEC7A, TLR7, and CCL3. As a prediction assuming that these tissues are actively growing, many of the 77 major cell types include states that progress from precursors to one or more terminally differentiated cell types. For example, cortical excitatory neurons show a continuous trajectory from PAX6+ neural precursors to NEUROD6+ differentiated neurons and further to SLC17A7+ mature neurons. In the liver, hepatic precursors (DLK1+, KRT8+, KRT18+) show a continuous trajectory to functional hepatoblasts (SLC22A25+, ACSS2+, ASS1+). In contrast to mouse organogenesis, where the maturation of transcriptional programs is tightly linked to developmental time, cell state trajectories correlated consistently with the estimated gestational age in these human data. The simplest explanation is that gene expression is significantly more dynamic during the early stages of development (i.e., organogenesis vs. fetal development). However, the heterogeneous representation and inaccuracies in estimated gestational age can also confound our elucidation.
[0388] In addition to the manual annotation of these cell types, Garnett was used to generate semi - automatic classifiers for each organ, as well as a global classifier. The Garnett classifier was generated without relying on clustering, using marker genes compiled individually from the literature. Classification by Garnett was highly consistent with manual classification. For example, 88% of the cells were in agreement in the pancreas (cluster expansion; 5% disagreed; 7% unclassified). Using the Garnett model trained on this human cell atlas, it was also possible to accurately classify cell types from other single - cell data sets, such as data from different methods and data from adult organs. For example, the inventors applied the Garnett classifier for the pancreas to inDrop single - cell RNA - seq data and found that this model accurately annotated 82% of the cells (cluster expansion; 11% inaccurate, 8% unclassified). These Garnett models have been posted on the inventors' website and can be widely used for the automatic classification of single - cell data from diverse organs.
[0389] Integration across tissues and investigation of unexpected cell types
[0390] Next, an attempt was made to integrate data across all 15 organs and compare cell types. To mitigate the effect of net differences in the number of sampled cells per organ and / or cell type, 5,000 cells per cell type were randomly sampled per organ (or, if fewer than 5,000 cells of a given cell type were shown in a given organ, all cells were obtained), and UMAP visualization was performed based on the genes that were most differentially expressed across cell types within each organ. As expected, cell types that were present in multiple organs, such as stromal cells, lymphatic endothelial cells, and mesodermal cells, were generally clustered together. For example, cell types related to development, such as diverse blood cells, PNS neurons, and mesenchyme, were also generally co - localized.
[0391] By leveraging this global UMAP, cell types that were not clearly annotated in organs not initially observed or were unanticipated were revealed. In many cases, co-localization with cell types annotated by global UMAP clarified their identity. For example, observing cells in the lung and adrenal gland that highly correlate with trophoblast giant cells from the placenta (e.g., expressing high levels of placental lactogen, chorionic gonadotropin, and aromatase) suggests that these are trophoblast cells (CSH1_CSH2 positive cells) that entered the fetal circulation. More surprisingly, observe cells in the placenta and spleen (AFP_ALB_positive cells) that highly correlate with hepatoblasts (e.g., expressing high levels of serum albumin, alpha-fetoprotein, and apolipoprotein).
[0392] In the heart, three cell types not predicted based on previous atlas-making efforts were observed. The first of these (SATB2_LRRC7 positive neurons) strongly correlates with CNS excitatory neurons and expresses markers including SATB2, PTPRD, and DAB1. As far as is known, this is an unexpected observation. While complete exclusion of contamination from other tissues cannot be achieved, these cells were observed in a consistent proportion (range) in each sampled heart (n = 9), and furthermore, no other CNS-like cell types were observed within the heart. The other two highly correlate with cardiomyocytes but express distinct programs that may reflect specialized roles. Specifically, ELF3_AGBL2 positive cardiomyocyte-like cells specifically express many genes associated with alveolar surfactant-secreting cells such as pulmonary surfactant protein 1 (SCGB3A2), surfactant-associated protein B (SFTPB), and surfactant-associated protein C (SFTPC), and CLC_IL5RA positive cardiomyocyte-like cells specifically express immune cell-related receptors such as interleukin 5 receptor subunit alpha (IL5RA) and hematopoietic-specific transmembrane protein 4 (MS4A3).
[0393] Characterization of cell type-specific gene regulatory networks and pathways.
[0394] Next, we examined the cell-type specific expression of surface and secreted protein-coding genes that are important for regulating cell-cell or cell-environment interactions. Most surface proteins (4,565 out of 5,480) and most secreted proteins (2,491 out of 2,933) were differentially expressed across 77 major cell types (FDR 0.05). For example, microglia specifically expressed both sialic acid-binding immunoglobulin-like lectin 8 (SIGLEC8) and oxidized LDL endocytosis receptor (OLR1), both of which are associated with Alzheimer's disease, and endothelial cells expressed both roundabout guidance receptor 4 (ROBO4) and endothelial cell adhesion molecule (ESAM), both of which are involved in angiogenesis and vascular patterning. Similarly, different neurons were labeled by distinct cell surface transporters. For example, in the cerebellum, specific expression of the glycine neurotransmitter transporter SLC6A5 in inhibitory interneurons, the excitatory amino acid transporter SLC1A6 in Purkinje neurons, the potassium channel KCNK9 in granule neurons, and the sodium / potassium / calcium exchanger SLC24A4 in SLC24A4_PEX5L-positive inhibitory interneurons was observed. There are numerous similar examples of cell-type specific expression of secreted proteins. Particularly interesting examples are the glycoprotein STC2, which is associated with all mesenchymal precursors or stem cells, and an unexpected cell type in the spleen (STC2_TLX1-positive cells) that specifically expresses TF TLX1 and NKX2-3.
[0395] Non-coding RNAs have been demonstrated to play important roles in normal development and disease. In these data, 3,130 out of 10,695 non-coding RNAs were differentially expressed across 77 major cell types (FDR 0.05). For example, ncRNAs were highly specific to microglia (RP11-489O18.1, RP11-480C22.1, RP11-10H3.1) or endothelial cells (AC011526.1, RP11-554D15.1, CTD-3179P9.1). The biological significance of such cell-type specific ncRNAs is unknown, but it is noteworthy that the pattern of their expression was sufficient to separate the 77 major cell types into developmentally coherent groups.
[0396] Most transcription factors (TFs) were also differentially expressed across 77 major cell types (1,715 out of 1,984, FDR 0.05). Many of the most specific TFs for each cell type were as expected, for example, RBPJL in acinar cells, OLG1 and OLG2 in oligodendrocytes, and PAX7 in satellite cells. In other cases, cell type-specific TFs were pointed out to consider unexpected cell types, such as stromal cell types characterized by the expression of lymphoid chemokines (CCL19_CCL21-positive cells) that specifically express TFs related to immune activation, as observed in the pancreas.
[0397] The inventors attempted to directly predict the interactions of TF target genes via gene expression data. Briefly, candidate interactions were identified by the covariance between TF expression and target gene expression across the complete dataset. These interactions were further filtered by ChIP-seq binding and motif enrichment analysis ("Methods"). 56,272 candidate TF-target gene links remained, including 706 TFs and 12,868 target genes. Of these 706 TF-bound gene sets, 220 showed enrichment of the corresponding TFs (FDR 0.05) in the manually clustered databases of the TF network (TRRUST) or Enrichr TF gene networks (for example, the most enriched TRRUST TF for the 330 genes binding to E2F1 is E2F1, adjusted p-value = 2.2e-14; the top Enrichr TF for the 1,219 genes binding to FLI1 is FLI1, adjusted p-value = 5.6e-122). When the target genes assigned to these 706 TFs were sorted and the analysis was repeated, none of the TF-bound gene sets were significantly enriched for the corresponding TFs at the same threshold.
[0398] Characterization of the development of the blood system across organs
[0399] The nature of this dataset provides an opportunity to investigate organ-specific differences in gene expression within widely occurring cell types, such as blood cells, endothelial cells, and epithelial cells. As a first such analysis, the inventors reclustered 103,766 cells from all organs corresponding to hematopoietic cell types. Subsequently, Louvain clustering and further annotation of fine-grained immune cell types were performed based on published gene markers. In some cases, very rare cell types were identified. For example, myeloid cells were separated into microglia, macrophages, and diverse dendritic cell subtypes (CD1C+, S100A9+, CLEC9A+, and pDC). The microglia cluster mainly originated from the cerebrum and cerebellum and was well separated from macrophages that corresponded to their different developmental origins. Lymphoid cells were clustered into several groups, including B cells, NK cells, ILC 3 cells, and T cells (the latter including thymic production trajectories). Also, very rare cell types such as plasma cells (139 cells, which are 0.1% of all blood cells or 0.003% of the complete dataset, mostly within the placenta) and TRAF1+ APCs (189 cells, which are 0.2% of all blood cells or 0.005% of the complete dataset, mostly within the thymus and heart) were recovered.
[0400] Gene expression markers for different immune cell types have been widely studied, but these can be limited by definitions through a restricted set of organs or cell types. In fact, the inventors found that many conventional immune cell markers are expressed in multiple cell types. For example, a conventional marker for T cells was also expressed in macrophages and dendritic cells (CD4) or NK cells (CD8A), in agreement with other studies. The inventors calculated pan-organ cell type-specific markers across 14 blood cell types. For example, T cells specifically expressed CD8B and CD5 as expected, but also expressed TENM1. ILC 3 cells annotated based on the expression of RORC and KIT were more specifically labeled by SORCS1 and JMY. These and other pan-organ defining markers may be useful for the labeling and purification of human fetal blood cell types in future studies.
[0401] As expected, different organs showed very different proportions of blood cells. For example, the liver contained the highest proportion of erythroblasts, consistent with its role as a major site of fetal erythrocytes, and T cells were concentrated in the thymus and B cells within the spleen. Blood cells recovered from the cerebellum and cerebrum were almost entirely microglia. Aggregate analysis also enabled the identification of rare cell populations in specific organs. For example, the inventors identified rare HSCs in the liver, spleen, and thymus, but also in the heart, lung, adrenal gland, and intestine.
[0402] Focusing on erythropoiesis, a continuous trajectory from HSCs to an intermediate cell type, erythroid-basophil-megakaryocyte biased progenitor cells (EBMPs) was observed, which then split into erythroid, basophil, and megakaryocyte trajectories, consistent with recent studies of the mouse fetal liver. This was consistent despite differences in species (human vs. mouse), technology (sci-RNA-seq3 vs. 10x), and organ (pan-organ vs. fetal organ). Unsupervised clustering was performed, adopting terminology from that study, and the continuum of the erythroid state was further divided into three stages, namely, early erythroid progenitor cells (EEP; labeled by SLC16A9 and FAM178B), committed erythroid progenitor cells (CEP; labeled by KIF18B and KIF15), and cells in the final differentiated state of erythrocytes (ETD; labeled by TMCC2 and HBB). The early and late stages of megakaryocytes were also readily identified. The corresponding dynamics of genome-wide chromatin accessibility in the erythroid lineage are further considered in the handbook.
[0403] Given the roles established for fetal erythrocytes as expected, a significant proportion of immune cells in the liver and spleen corresponded to EEP, CP, and megakaryocyte progenitor cells. Surprisingly, EEP, CEP, and megakaryocyte progenitor cells were also observed in the adrenal gland in the nuclear samples studied. Since more common cell types were not observed in the liver and spleen, species contamination during adrenal gland recovery cannot be said to account for the findings. Confirmation by orthogonal methods is required, but the results suggest the possibility of the adrenal gland as an additional site of fetal erythrocytes.
[0404] Macrophages are even more widely distributed. Next, all macrophages together with microglia from the brain were stained and independently subjected to UMAP visualization and Louvain clustering. Microglia were divided into three subclusters, and one of them labeled with IL1B and TNFRSF10D is likely to represent activated microglia involved in the inflammatory response. The other microglia clusters were labeled by the expression of TMEM119 and CX3CR1 (more common in the cerebrum) or PTPRC and CDC14B (more common in the cerebellum).
[0405] Macrophages outside the brain were clustered into three major groups, namely, 1) antigen-presenting macrophages, mostly found in the GI tract organs (intestine and stomach), labeled by the high expression of antigen presentation (HLA DPB1, HLA DQA1) and inflammatory activation (AHR) genes, 2) perivascular macrophages, found in most organs, with specific expression of markers such as F13A1 and COLEC12, as well as novel markers such as RNASE1 and LYVE1, and 3) phagocytic macrophages, concentrated in the liver, spleen, and adrenal glands, with specific expression of markers such as CD5L, TIMD4, and VCAM1. Phagocytic macrophages are important for erythrophagocytosis, and these observations in the adrenal glands are consistent with the aforementioned potential role as a site of fetal erythropoiesis.
[0406] Characterization of endothelial and epithelial cells across organs
[0407] As a second analysis of single cell types across many organs, the inventors reclustered cells from all organs corresponding to vascular endothelium, lymphatic endothelium, or endocardium. These three groups were easily separable from each other, and vascular endothelial cells were further clustered to some extent for each organ. The organ-specific differences were more easily detected than the differences between arteries, capillaries, and veins, and were consistent with previous cell atlases of adult mice.
[0408] Differential gene expression analysis identified 700 markers that are specifically expressed in subsets of endothelial cells (FDR 0.05, >2-fold difference in expression between the top and second clusters). For approximately one-third (236 out of 700) of these coding membrane proteins, many appeared to correspond to potential specialized functions. For example, renal endothelial cells specifically expressed acid-sensing ion channel 2 (ASIC2), a mechanosensor involved in myogenic contraction and blood flow regulation within the kidney. Lung endothelial cells specifically expressed relaxin family peptide receptor 1 (RXFP1). RXFP1 specifically expressed the sodium-dependent lysophosphatidylcholine cotransporter 1 (MFSD2A), which is involved in endogenous nitric oxide-mediated vasorelaxation in the lung, and MFSD2A is integrally involved in the establishment and function of the blood-brain barrier. Potential regulatory criteria for differential gene expression in subsets of endothelium are discussed in the manual.
[0409] As a third analysis of widely distributed cell types, epithelial cells from all organs were reclustered and subjected to UMAP visualization. Some epithelial cell types are highly organ-specific; for example, adenocarcinoma (pancreas) and alveolar cells (lung), and epithelial cells with similar functions generally cluster together. For example, the expression programs of squamous epithelial cells (lung, stomach) co-cluster with corneal and conjunctival epithelial cells (eye), and PDE1C_ACSM3-positive cells (stomach) co-cluster with intestinal epithelial cells (intestine).
[0410] Within epithelial cells, two neuroendocrine cell clusters were identified. The simpler of these corresponded to adrenal chromaffin cells and were labeled by the specific expression of HHMX1 (NKX-5-3), a TF involved in the diversification of sympathetic neurons. The other cluster included neuroendocrine cells from multiple organs (stomach, intestine, pancreas, lung) and was labeled by the specific expression of NKX2-2, a TF that plays an important role in islet and enteroendocrine differentiation. The inventors performed further analysis on the latter group and identified five subsets: 1) pancreatic islet β cells labeled by insulin expression, 2) pancreatic islet α / γ cells labeled by pancreatic polypeptide expression and glucagon expression, 3) pancreatic islet δ cells labeled by somatostatin expression, 4) pulmonary neuroendocrine cells (PNEC) labeled by the expression of ASCL1, a TF that plays an important role in identifying this lineage in the lung, and 5) enteroendocrine cells. Enteroendocrine cells further included multiple subsets such as NEUROG-expressing pancreatic islet ε progenitor cells, TPH1-expressing chromaffin cells both in the stomach and intestine, gastrin-expressing or cholecystokinin-expressing G / L / K / I cells. Finally, ghrelin-expressing enteroendocrine progenitor cells in the stomach and intestine were observed, but ghrelin-expressing endocrine cells in the developing lung were also observed. Since the diverse functions of neuroendocrine cells are closely associated with their secreted proteins, 1,086 secreted protein-coding genes that are differentially expressed across neuroendocrine cells were identified (FDR 0.05). For example, PNEC showed specific expression of trefoil factor 3, which is involved in mucosal protection and lung calcified cell differentiation, gastrin-releasing peptide, which stimulates gastrin release from G cells in the stomach, and SCGB3A2, a surfactant associated with lung development.
[0411] As an illustrative example of a method for exploring cell trajectories using these data, we further investigated the pathway of diversification of epithelial cells leading to renal tubular cells. We combined and reclustered ureteric bud metanephric cells to identify both progenitor cells and terminal renal epithelial cell types, and the differentiation pathway was highly consistent with recent studies of the human fetal kidney. Differential gene expression analysis was used to further evaluate the properties of TFs that potentially regulate their specification. For example, nephron progenitor cells in the metanephric trajectory expressed high levels of mesenchymal and Meis homeobox genes (MEOX1, MEIS1, MEIS2), and podocytes specifically expressed MAFB and TCF21 / POD1. As another example, HNF4A was specifically expressed in proximal tubular cells, and mutations in this gene cause Fanconi syndrome, a disease that specifically affects the proximal tubules. This has recently been shown to be required for the formation of proximal tubules in mice.
[0412] Comparison of human and mouse developmental atlases
[0413] To investigate the developmental relationships between cell types, we next compared these data with our recent mouse organogenesis cell atlas (MOCA), which profiled 2 million cells from whole embryos spanning the earlier mammalian developmental window of E9.5 - E13.5.
[0414] As a first approach, the 77 major human cell types defined herein were compared to the developmental trajectories defined by MOCA by the cell type cross-matching method described above. Briefly, this method uses non-negative least squares (NNLS) regression to select the cell type pairs that best match each other from two datasets. Most human cell types strongly match a single major mouse trajectory and sub-trajectories. These generally correspond to expected values and serve as a form of validation for both sets of annotations. Some discrepancies facilitated important corrections to the MOCA annotations. Many of the human cell types and mouse trajectories that lacked a strong match (summed NNLS regression coefficient <0.6) corresponded to tissues excluded in other datasets (e.g., mouse placenta, human skin, and gonads). Other ambiguities are probably due to gaps between the developmental windows studied (e.g., adrenal cell types), rarity (e.g., bipolar cells), and / or complex relationships between cell types (e.g., fetal cell types derived from multiple embryonic trajectories).
[0415] As a second approach, we attempted to cluster human and mouse cells together. Briefly, we sampled 100,000 mouse embryonic cells (random) and 65,000 human fetal cells (up to 1,000 cells from each of 77 cell types) from MOCA and subjected them to the recently described Seurat strategy to integrate the cross-species scRNA-seq dataset. The distribution of mouse cells in the resulting UMAP-based visualization was very similar to our global analysis of MOCA. Furthermore, surprisingly, the cells were distributed in a generally reasonable manner with respect to both developmental and temporal relationships, rather than spatial organ location. For example, we observed that human fetal endothelial cells, hematopoietic cells, hepatocytes, epithelial cells, and mesenchymal cells were all mapped to the corresponding mouse embryonic trajectories. Human fetal brain neurons and cerebellar neurons overlapped with the mouse embryonic neural tube trajectories, but, perhaps due to excessive differences between species or developmental stages, human fetal neural crest derivatives, such as ENS neurons, visceral neurons, sympathetic neuroblasts, and chromaffin cells, were clustered separately from the corresponding mouse embryonic trajectories. As expected, human ENS glia as well as Schwann cells overlapped with the mouse embryonic PNS glial trajectories. Human fetal astrocytes were clustered with the mouse embryonic neuroepithelial trajectories (mouse astrocytes do not develop until E18.5). Human fetal oligodendrocytes overlapped with a rare mouse embryonic sub-trajectory (Pdgfra+ glia) that, upon reflection, corresponds to oligodendrocyte progenitor cells (OPCs; Olig1+, Olig2+, Brinp3+), casting doubt on the previous annotation of distinct Oligo1+ sub-trajectories as oligodendrocyte precursors.
[0416] To visualize the more detailed relationships between human fetal cells and mouse embryonic cells, a similar integrative analysis strategy was applied to extract human and mouse cells from hematopoietic, endothelial, and epithelial trajectories. Data from this fetal human cell atlas enables the "whole embryo" mouse data to be readily deconvolved into granulated functional or spatial groups. For example, subsets of the mouse "leukocyte" trajectory map to specific human blood cell types such as HSC, microglia, macrophages (liver and spleen), macrophages (other organs), and DC. These subsets were further evidenced by the expression of relevant blood cell markers. Similarly, the inventors observed that related subsets of mouse / human endothelial and epithelial cells map to each other. This approach may be useful for obtaining gene expression programs of progenitor cells of specific lineages at developmental time points where access or anatomical dissection is difficult. For example, within mouse cells previously labeled as the foregut epithelial trajectory, it is possible to deconstruct factors likely to be activated in the stomach versus the pancreas.
[0417] Discussion
[0418] The successful development of a functional human fetus is a remarkable process, characterized by the processes of cell proliferation and differentiation over three major developmental stages.
[0419] Following a short (2 weeks from fertilization) embryonic period with simple cell proliferation and implantation in the uterus, the embryogenesis stage continues gastrulation, neurulation, and organogenesis, characterized by intense cell differentiation and the generation of visceral organ precursors. By the end of the 10th week of gestation, the embryo has acquired a basic form called a fetus. Over the next 20 weeks, various organs continue to grow and mature, and diverse terminally differentiated cell types are generated from precursors.
[0420] Both the embryonic and embryogenic stages have been intensively profiled at single-cell resolution in humans or model systems (i.e., mice) using a shared early developmental program. The late developmental stage (fetal stage) exhibits different developmental programs and durations in Homo sapiens compared to other species. Also, because organs are more complex and there are technical limitations, it is difficult to obtain a complete picture of cell dynamics at this stage. Although several recent studies on single cells in fetal development have been published, most of these are limited to specific organs or cell lineages and do not provide a complete picture of organ-wide development.
[0421] Materials and Methods:
[0422] Culture and Nuclear Extraction of Mammalian Cells
[0423] All mammalian cells were cultured at 5% CO2, 37 °C and maintained in high-glucose DMEM (Gibco catalog number 11965) supplemented with 10% FBS and 1X Pen / Strep (Gibco catalog number 15140122; 100 U / mL penicillin, 100 μg / mL streptomycin). Cells were trypsinized with 0.25% trypsin-EDTA (Gibco catalog number 25200-056) and split 1:10 three times a week.
[0424] All cell lines were trypsinized, spun down at 300 x g for 5 minutes (4°C), and washed once with 1X ice-cold PBS. 5M cells were combined and lysed using 1 mL of ice-cold cell lysis buffer (modified to contain 10 mM Tris-HCl, pH 7.4, 10 mM NaCl, 3 mM MgCl2, and 0.1% IGEPAL CA-630, 1% SUPERase In RNase inhibitor, and 1% BSA). The filtered nuclei were then transferred to a new 15 mL tube (Falcon), pelleted by centrifugation at 500 x g, 4°C for 5 minutes, and washed once with 1 mL of ice-cold cell lysis buffer. The nuclei were fixed in 4 mL of ice-cold 4% paraformaldehyde (EMS) for 15 minutes on ice. After fixation, the nuclei were washed twice in 1 mL of nuclear wash buffer (cell lysis buffer without IGEPAL) and resuspended in 500 uL of nuclear wash buffer. 100 uL of the sample was placed in each tube, divided into 5 tubes, and snap-frozen in liquid nitrogen.
[0425] Preparation and nuclear extraction of human fetal tissues
[0426] Human fetal tissues were combined and processed to reduce batch effects. Each organ was pulverized to tissue powder with a hammer (on dry ice) and mixed prior to sampling. First, 1 mL of ice-cold cell lysis buffer (10 mM Tris-HCl, pH 7.4, 10 mM NaCl, 3 mM MgCl2, and 0.1% IGEPAL CA-630 53 , 1% SUPERase Powders of 0.1 - 1 g were incubated with In and modified to also contain 1% BSA, and then transferred onto a 40 - μm cell strainer (Falcon). The tissue was homogenized using the rubber tip of a syringe plunger (5 mL, BD) in 4 mL of cell lysis buffer. Next, the filtered nuclei were transferred to a new 15 - mL tube (Falcon), pelleted by centrifugation at 500 x g for 5 minutes, and washed once with 1 mL of cell lysis buffer. The nuclei were fixed in 5 mL of ice - cold 4% paraformaldehyde (EMS) on ice for 15 minutes. After fixation, the nuclei were washed twice in 1 mL of nuclear wash buffer (cell lysis buffer without IGEPAL) and resuspended in 500 μL of nuclear wash buffer. 250 μL of the sample was placed in each tube, divided into two tubes, and snap - frozen in liquid nitrogen. For human cell extraction and paraformaldehyde fixation in some organs (kidney, pancreas, intestine, and stomach).
[0427] Preparation and sequencing of sci - RNA - seq3 libraries
[0428] Similar to the published sci-RNA-seq3 protocol, paraformaldehyde-fixed nuclei were processed with minor modifications. Briefly, thawed nuclei were permeabilized on ice for 3 minutes with 0.2% TritonX-100 (in nuclear wash buffer), followed by brief sonication (12 seconds in low-power mode on Diagenode) to reduce nuclear aggregation. The nuclei were then washed once with nuclear wash buffer and filtered through a 1 mL Flowmi cell strainer (Flowmi). The filtered nuclei were spun down at 500 x g for 5 minutes and resuspended in nuclear wash buffer. Nuclei from each sample were then dispensed into multiple individual wells in four 96-well plates. The link between well ID and mouse embryo was recorded for downstream data processing. For each well, 80,000 nuclei (16 μL) were mixed with 8 μL of 25 μM immobilized oligo-dT primer ((5’- / 5Phos / CAGAGCNNNNNNNN[10bp barcode]TTTTTTTTTTTTTTTTTTTTTTTTTTTTTT-3’ (SEQ ID NO: 1), where "N" is any base in the sequence; IDT) and 2 μL of 10 mM dNTP mix (Thermo), denatured at 55 °C for 5 minutes, and immediately placed on ice. Then, 14 μL of first-strand reaction mix containing 8 μL of 5X Superscript IV First-Strand Buffer (Invitrogen), 2 μL of 100 mM DTT (Invitrogen), 2 μL of SuperScript IV reverse transcriptase (200 U / μL, Invitrogen), and 2 μL of RNaseOUT Recombinant Ribonuclease Inhibitor (Invitrogen) was added to each well. Reverse transcription was performed by incubating the plate at gradient temperatures (2 minutes at 4 °C, 2 minutes at 10 °C, 2 minutes at 20 °C, 2 minutes at 30 °C, 2 minutes at 40 °C, 2 minutes at 50 °C, and 10 minutes at 55 °C).
[0429] After the reverse transcription reaction, 60 μL of nuclear dilution buffer (10 mM Tris-HCl, pH 7.4, 10 mM NaCl, 3 mM MgCl2, and 1% BSA) was added to each well. The nuclei from all wells were pooled together and spun down at 500 x g for 10 minutes. Then, the nuclei were resuspended in nuclear wash buffer, and 20 μL of Quick Ligase buffer (NEB), 2 μL of Quick DNA Ligase (NEB), 10 μL of nuclei in nuclear wash buffer, and 8 μL of barcoded ligation adapter (100 μM, 5’-GCTCTG[9bp or 10bp barcode A] / dideoxyU / ACGACGCTCTTCCGATCT[reverse complement of barcode A]-3’ (SEQ ID NO: 2)) were redistributed into another four 96-well plates containing each well. The ligation reaction was carried out at 25°C for 10 minutes. After the ligation reaction, 60 μL of nuclear dilution buffer (10 mM Tris-HCl, pH 7.4, 10 mM NaCl, 3 mM MgCl2, and 1% BSA) was added to each well. The nuclei from all wells were pooled together and spun down at 600 x g for 10 minutes.
[0430] The nuclei were washed once with nuclear wash buffer, filtered once through a 1 mL Flowmi cell strainer (Flowmi), counted, and distributed into eight 96-well plates containing 2,500 each in 5 μL of nuclear wash buffer per well and 3 μL of elution buffer (Qiagen). Then, 1.33 μL of mRNA second strand synthesis buffer (NEB) and 0.66 μL of mRNA second strand synthesis enzyme (NEB) were added to each well, and second strand synthesis was carried out at 16°C for 180 minutes.
[0431] For tagging, each well was mixed with 11 μL of Nextera TD buffer (Illumina) and 1 μL of i7-only TDE1 enzyme (62.5 nM, diluted with Nextera TD buffer (Illumina)), and then incubated at 55 °C for 5 minutes for tagging. Next, the reaction was stopped by adding 24 μL of DNA binding buffer (Zymo) per well and incubated at room temperature for 5 minutes. Then, each well was purified using 1.5x AMPure XP beads (Beckman Coulter). In the elution step, 8 μL of nuclease-free water, 1 μL of 10X USER buffer (NEB), and 1 μL of USER enzyme (NEB) were added to each well and incubated at 37 °C for 15 minutes. Another 6.5 μL of elution buffer was added to each well. The AMPure XP beads were removed by a magnetic stand, and the elution product (16 μL) was transferred to a new 96-well plate.
[0432] For PCR amplification, each well (16 μL of product) was mixed with 2 μL of 10 μM indexed P5 primer (5’-AATGATACGGCGACCACCGAGATCTACAC[i5]ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3’ (SEQ ID NO: 3); IDT), 2 μL of 10 μM P7 primer (5’-CAAGCAGAAGACGGCATACGAGAT[i7]GTCTCGTGGGCTCGG-3’ (SEQ ID NO: 4), IDT), and 20 μL of NEBNext High-Fidelity 2x PCR MASTER Mix (NEB). Amplification was performed using a program of 5 minutes at 72 °C, 30 seconds at 98 °C, 12 - 16 cycles of “10 seconds at 98 °C, 30 seconds at 66 °C, 1 minute at 72 °C”, and finally 5 minutes at 72 °C.
[0433] After PCR, the samples were pooled and purified using 0.8 volume of AMPure XP beads. Library concentration was determined by Qubit (Invitrogen), and the library was visualized by electrophoresis on a 6% TBE-PAGE gel. All libraries were sequenced on one NovaSeq platform (Illumina) (Read 1: 34 cycles, Read 2: 52 cycles, Index 1: 10 cycles, Index 2: 10 cycles).
[0434] For paraformaldehyde-fixed cells, with slight modifications, they were processed as follows in the same manner as the fixed nuclei. That is, the frozen-fixed cells were thawed in a 37 °C water bath, spun down at 500 xg for 5 minutes, and incubated on ice for 3 minutes with 500 μL of PBSR (1x PBS, pH 7.4, 1% BSA, 1% SuperRnaseIn, 1% 10 mM DTT). The cells were pelleted and resuspended in 500 μL of nuclease-free water containing 1% SuperRnaseIn. To incubate on ice for 5 minutes, 3 mL of 0.1 N HCl was added to the cells (7). To neutralize the HCl, 3.5 mL of Tris-HCl (pH = 8.0) a...
Claims
【Claim 1】 The invention described in the specification.
Citation Information
Patent Citations
High-throughput single-cell sequencing with reduced amplification bias
WO2019222688A1