High-throughput single-cell library, as well as method for manufacturing and using it.

The transpososome complex method addresses the challenge of characterizing rare cells by generating indexed nucleic acids without custom transposons, enabling efficient construction of sequencing libraries and identification of cell subpopulations.

JP7864297B2Active Publication Date: 2026-05-25ILLUMINA INC +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
ILLUMINA INC
Filing Date
2020-12-18
Publication Date
2026-05-25

AI Technical Summary

Technical Problem

Current single-cell genomics technologies face challenges in comprehensively characterizing rare cells within a population without enrichment, which is costly and difficult, and often require custom-modified transposons for single-step labeling.

Method used

A method using a transpososome complex for single-cell combinatorial indexing that eliminates the need for custom-modified transposons, involving a transposase and universal sequences, with optional partitioning into compartments for indexing and pooling of nucleic acids.

Benefits of technology

Enables comprehensive characterization of rare cells by generating indexed, double-indexed, or triple-indexed nucleic acids, facilitating the construction of sequencing libraries and identification of subpopulations based on biological features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007864297000009
    Figure 0007864297000009
  • Figure 0007864297000010
    Figure 0007864297000010
  • Figure 0007864297000011
    Figure 0007864297000011
Patent Text Reader

Abstract

Provided herein is a method for preparing a sequencing library containing nucleic acids from a plurality of single cells. In one embodiment, the sequencing library contains nucleic acids representing chromatin accessibility from a plurality of single cells. In one embodiment, the nucleic acids include three index sequences. In another embodiment, the present disclosure provides a method for characterizing rare events in isolated cells and nuclei. In an embodiment, providing can include providing a plurality of nuclei or cells in a plurality of compartments, each compartment containing a subset of nuclei or cells or representing a sample.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Cross-reference of related applications)

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 950,670, filed on 19 December 2019, which is incorporated herein by reference in its entirety.

[0003] (Government investment)

[0004] This invention was made with government support under authorization number T32 HL007828 granted by the National Institutes of Health. The government has certain rights in this invention.

[0005] (Field of Invention)

[0006] Embodiments of this disclosure relate to nucleic acid sequencing. Specifically, embodiments of methods and compositions provided herein relate to preparing a single-cell combinatorial index sequencing library and obtaining sequence data therefrom. In some embodiments, the sequence data obtained from the library is comprehensive, and in other embodiments, the sequence data obtained from the library enables the characterization of rare events.

[0007] [Background technology]

[0008] Single-cell combinatorial indexing ("sci-") is a methodological framework for constructing single-cell combinatorial sequencing libraries by uniquely labeling the nucleic acid contents of multiple single cells or single nuclei using split-pool barcoding. Current single-cell genomics technologies often involve adding unique labels in a single step using transposomal complexes, but this requires a large number of custom-modified transposons.

[0009] Single-cell genomics techniques resolve cell-to-cell differences that are difficult to determine when studying bulk populations of cells. In many important applications such as oncology, immunology, and metagenomics, there is great interest in and challenges with characterizing rare cells. Current methods of single-cell sequencing can characterize millions of single cells in parallel. However, comprehensively characterizing rare cells within a population based on sequencing without enrichment is costly and difficult.

[0010] SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0011] Provided herein is a method of using a transpososome complex during single-cell combinatorial indexing without the need to generate custom-modified transposons.

[0012] In one embodiment, the present disclosure provides a method for preparing a sequencing library comprising nucleic acids from a plurality of single nuclei or single cells. The method includes providing a plurality of nuclei or cells, wherein the nuclei or cells contain nucleosomes, and contacting the plurality of nuclei or cells with a transposome complex comprising a transposase and a universal sequence. In one embodiment, the plurality of nuclei or cells are in bulk upon contact with the transposome complex, and in another embodiment, upon contact with the transposome complex, the plurality of nuclei or cells are partitioned within a first plurality of compartments, each compartment containing a subset of the nuclei or cells or representing a sample. Contacting further includes conditions suitable for incorporating the universal sequence into the DNA nucleic acid, resulting in a double-stranded DNA nucleic acid containing the universal sequence. In embodiments where contacting occurs with the plurality of nuclei or cells in bulk, the method also includes partitioning the plurality of nuclei or cells into a first plurality of compartments, each compartment containing a subset of the nuclei or cells. The DNA molecules within each subset of nuclei or cells are processed to generate indexed nuclei or cells. This processing includes adding a first compartment-specific index sequence to the DNA nucleic acids present in each subset of nuclei or cells, resulting in indexed nucleic acids present in the indexed nuclei or cells. This processing can include ligation, primer extension, hybridization, amplification, or combinations thereof. The indexed nuclei or cells can be combined to generate pooled indexed nuclei or cells.

[0013] In one embodiment, providing can include providing the plurality of nuclei or cells in a plurality of compartments, each compartment containing a subset of the nuclei or cells or representing a sample. Contacting can include contacting each compartment with the transposome complex, and the method can further include combining the nuclei or cells after contact to generate pooled nuclei or cells.

[0014] In one embodiment, contacting comprises contacting each subset with two transposome complexes, one of which comprises a first transposase containing a first universal sequence, and the second transposome complex comprises a second transposase containing a second universal sequence, and the contacting further comprises conditions suitable for incorporating the first and second universal sequences into a DNA nucleic acid to yield a double-stranded DNA nucleic acid containing the first and second universal sequences.

[0015] In one embodiment, the method may further include distributing a pool of indexed nuclei or cells, each containing an indexed nucleus or cell, into a second set of compartments, each compartment containing a subset of nuclei or cells, and processing the DNA molecules within each subset of nuclei or cells to generate double-indexed nuclei or cells. Processing may include adding a second compartment-specific index sequence to the DNA nucleic acids present in each subset of nuclei or cells to result in double-indexed nucleic acids present in the indexed nuclei or cells. The method may also include combining the double-indexed nuclei or cells to generate a pool of double-indexed nuclei or cells.

[0016] In one embodiment, the method may further include distributing a pool of indexed nuclei or cells, including double-indexed nuclei or cells, into a third number of compartments, each compartment containing a subset of nuclei or cells, and processing the DNA molecules within each subset of nuclei or cells to generate triple-indexed nuclei or cells. Processing may include adding a third compartment-specific index sequence to the DNA nucleic acids present in each subset of nuclei or cells to result in triple-indexed nucleic acids present in the indexed nuclei or cells. The method may also include combining triple-indexed nuclei or cells to generate a pool of triple-indexed nuclei or cells.

[0017] In one embodiment, the method may further include obtaining indexed nucleic acids (e.g., double-indexed, triple-indexed, etc.) from a pool of indexed nuclei or cells, and thus constructing a sequencing library from multiple nuclei or cells.

[0018] Furthermore, this specification provides methods for identifying and / or characterizing subpopulations of cells. In one embodiment, the method includes providing a sequencing library, such as a single-cell combinatorial sequencing library. Optionally, the sequencing library is constructed from a population of cells or nuclei enriched with a particular characteristic. The method may include scrutinizing the sequencing library by targeted sequencing. Targeted sequencing may be based on biological features typically present in a small percentage of the cells used to construct the library. Examples of biological features include, but are not limited to, nucleotide sequences indicating cell class, species type, or disease state. In addition to targeted sequencing of biological features, sequencing also includes determining the sequence of index sequences present in the same modified target nucleic acids as the biological features. As a result, members of the sequencing library that originate from the same cells or nuclei as the library members containing the biological features are identified. The method further includes modifying the sequencing library to increase the expression of these members that originate from the same cells or nuclei as the library members containing the biological features. This modification may include enriching the sequencing library with desired members or depleting the sequencing library with undesirable members to obtain a sublibrary.

[0019] definition

[0020] Unless otherwise specified, terms used herein will be understood to have their common meanings in the relevant art. Some terms used herein and their meanings are listed below.

[0021] As used herein, the terms “living organism” and “subject” are interchangeable and refer to microorganisms (e.g., prokaryotes or eukaryotes), animals, and plants. Examples of animals include mammals such as humans.

[0022] As used herein, the term “cell type” is intended to identify a cell based on morphology, phenotype, developmental origin, or other known or recognizable distinguishable cellular characteristics. A variety of different cell types can be obtained from a single organism (or from organisms of the same species). Exemplary cell types include gametes (e.g., female gametes such as oocytes or egg cells, and male gametes such as sperm), ovarian epithelium, ovarian fibroblasts, testes, bladder, immune cells, B cells, T cells, natural killer cells, dendritic cells, cancer cells, eukaryotic cells, stem cells, blood cells, muscle cells, adipocytes, skin cells, nerve cells, osteocytes, pancreatic cells, endothelial cells, pancreatic β cells, pancreatic endothelium, bone marrow lymphoblasts, bone marrow macrophages, myeloblasts, bone marrow adipocytes, bone marrow osteoblasts, bone marrow chondrocytes, bone marrow chondrocytes, bone marrow chondrocytes, myeloblasts, bone marrow chondrocytes, promyeloblasts, bone marrow megakaryoblasts, bladder, brain B lymphocytes, brain glial cells, neurons, brain astrocytes, neuroectoderm, brain macrophages, brain microglia, brain epithelium, cortical neurons, brain fibroblasts, mammary epithelium, colon epithelium, colon B lymphocytes Pacocytes, mammary epithelium, mammary myoepithelium, mammary fibroblasts, enteric cells, cervical epithelium, mammary ductal epithelium, tongue epithelium, tonsillar dendritic cells, tonsillar lymphocytes, peripheral blood lymphoblasts, peripheral blood T lymphoblasts, peripheral blood T lymphocytes, peripheral blood natural killer cells, peripheral blood B lymphoblasts, peripheral blood monocytes, peripheral blood myeloblasts, peripheral blood monoblasts, peripheral blood monoblasts, peripheral blood monoblasts, peripheral blood monoblasts, peripheral blood T lymphocytes, peripheral blood Examples of such cells include, but are not limited to, peripheral blood premyeloblasts, peripheral blood macrophages, peripheral blood basophils, liver endothelium, liver mast, liver epithelium, liver B lymphocytes, spleen endothelium, spleen epithelium, spleen B lymphocytes, hepatocytes, liver cells, fibroblasts, lung epithelium, bronchial epithelium, lung fibroblasts, lung B lymphocytes, lung Schwann cells, lung squamous epithelium, lung macrophages, lung osteoblasts, neuroendocrine cells, lung alveoli, gastric epithelium, and gastric fibroblasts. In one embodiment, the various different cell types obtained from a single organism may include cells of the organism and other cells such as cells of symbiotic or pathogenic microorganisms associated with the organism. Examples of symbiotic or pathogenic microorganisms associated with an organism include, but are not limited to, prokaryotic and eukaryotic microorganisms present in or within tissues of a microbiome sample derived from the organism and which selectively cause disease.

[0023] As used herein, the term “tissue” is intended to mean a collection or aggregate of cells that work together to perform one or more specific functions within an organism. Cells may be morphologically similar by choice. Exemplary tissues include, but are not limited to, embryonic tissue, epididymis, eye, muscle, skin, tendon, vein, artery, blood, heart, spleen, lymph node, bone, bone marrow, lung, bronchi, trachea, intestine, small intestine, large intestine, colon, rectum, salivary gland, tongue, gallbladder, appendix, liver, pancreas, brain, stomach, skin, kidney, ureter, bladder, urethra, gonad, testicle, ovary, uterus, fallopian tube, thymus, pituitary gland, thyroid gland, adrenal gland, or parathyroid gland. Tissues may originate from any of the various organs of a human or other organism. Tissues may be healthy or unhealthy. Examples of unhealthy tissues include malignant tumors of reproductive tissue, lungs, breasts, colorectal cavity, prostate, nasopharynx, stomach, testes, skin, nervous system, bones, ovaries, liver, blood tissue, pancreas, uterus, kidneys, and lymphoid tissue. Malignant tumors may be, but are not limited to, various histological subtypes, such as carcinoma, adenocarcinoma, sarcoma, fibroadenocarcinoma, neuroendocrine, or undifferentiated.

[0024] As defined herein, “Sample” and its derivatives are used in their broadest sense and include any specimen, culture, etc., suspected to contain target nucleic acids and / or target proteins. In some embodiments, a sample includes DNA, RNA, proteins, or combinations thereof. A sample may include any biological, clinical, surgical, agricultural, atmospheric, or aquatic-based specimen containing one or more nucleic acids and / or one or more proteins. The term also includes any isolated nucleic acids from a sample, such as genomic DNA or transcriptome, and any isolated proteins from a sample. In some embodiments, a sample includes an aggregate of cells or nuclei.

[0025] As used herein, the term “compartment” is intended to mean an area or volume that separates or isolates something from other things. Exemplary compartments include, but are not limited to, vials, tubes, wells, droplets, boluses, beads, containers, surface features, or areas or volumes separated by physical forces such as fluid flow, magnetism, or electric current. In one embodiment, a compartment is a well in a multiwell plate, such as a 96 or 384-well plate. In one embodiment, a compartment is a well with a patterned surface (e.g., a microwell or nanowell). As used herein, a droplet may include a hydrogel bead, which is a bead for encapsulating one or more nuclei or cells and contains a hydrogel composition. In some embodiments, a droplet is a homogeneous droplet of hydrogel material or a hollow droplet having a polymer hydrogel shell. Whether homogeneous or hollow, a droplet may be capable of encapsulating one or more nuclei or cells. In some embodiments, a droplet is a surfactant-stabilized droplet.

[0026] As used herein, “transposome complex” refers to a nucleic acid comprising an embedded enzyme and an embedded recognition site. A “transposome complex” is a functional complex formed by a transposase and a transposase recognition site capable of catalyzing a rearrangement reaction (see, for example, Gunderson et al., International Publication 2016 / 130704). Examples of embedded enzymes include, but are not limited to, integrases or transpores. Examples of embedded recognition sites include, but are not limited to, transposase recognition sites.

[0027] As used herein, the term “nucleic acid” is used interchangeably with “polynucleotide” and “oligonucleotide.” Nucleic acid includes naturally occurring nucleic acids or their functional analogues, intended to be consistent with its use in the art. Particularly useful functional analogues can hybridize to nucleic acids in a sequence-specific manner or can be used as templates for replicating specific nucleotide sequences. Naturally occurring nucleic acids generally have a backbone containing a phosphodiester bond. Analog structures may have alternative backbone bonds containing any of the various known in the art. Naturally occurring nucleic acids generally have a deoxyribose sugar (e.g., found in deoxyribonucleic acid (DNA)) or a ribose sugar (e.g., found in ribonucleic acid (RNA)). Nucleic acid may contain any of the various analogues of these sugar moieties known in the art. Nucleic acid may contain natural or non-natural bases. In this regard, natural deoxyribonucleic acid may have one or more bases selected from the group consisting of adenine, thymine, cytosine, or guanine, and ribonucleic acid may have one or more bases selected from the group consisting of adenine, uracil, cytosine, or guanine. Useful non-natural bases that may be included in nucleic acids are known in the art. Examples of non-natural bases include locked nucleic acid (LNA), cross-linked nucleic acid (BNA), and pseudo-complementary bases (Trilink Biotechnologies, San Diego, California). LNA and BNA bases can be incorporated into DNA oligonucleotides to enhance the hybridization strength and specificity of the oligonucleotides. LNA and BNA bases, and the use of such bases, are known and commonplace to those skilled in the art. Unless otherwise stated, the term “nucleic acid” includes natural and non-natural DNA, mRNA, and non-coding RNA, e.g., RNA without poly(A) at the 3' end, and nucleic acids derived from RNA, e.g., cDNA. The term “nucleic acid” refers only to the primary structure of a molecule. Therefore, this term includes triple-stranded, double-stranded, and single-stranded deoxyribonucleic acid ("DNA"), as well as triple-stranded, double-stranded, and single-stranded ribonucleic acid ("RNA").

[0028] As used herein, the term “target” is intended to be a semantic identifier of a molecule whose source, function, identity, and / or composition are being investigated. Examples of targets include, but are not limited to, nucleic acids and proteins. As used herein, when the term “target” is used in relation to nucleic acids, it is intended to be a semantic identifier of the nucleic acid in the context of the methods or compositions described herein, and does not necessarily limit the structure or function of the nucleic acid other than those expressly indicated elsewhere. The target nucleic acid may be any nucleic acid of an essentially known or unknown sequence. This may be, for example, a fragment of genomic DNA (e.g., chromosomal DNA), extrachromosomal DNA such as plasmids, cell-free DNA, RNA (e.g., RNA or non-coding RNA), protein (e.g., cellular or cell surface protein), or cDNA. The target nucleic acid may be a nucleic acid that binds to a compound such as an antibody that specifically binds biomolecules such as proteins, glycans, proteoglycans, or lipids (U.S. Patent Application Publication 2018 / 0273933). Sequencing may result in the determination of the sequence of all or part of the target molecule. The target may originate from a primary nucleic acid sample such as a nucleus. In one embodiment, the target can be processed into a template suitable for amplification by placing a universal sequence at one or both ends of each target fragment. The target can also be obtained from a primary RNA sample by reverse transcription to cDNA. In one embodiment, the target is used by reference to a subset of DNA, RNA, or protein present in a cell. Target sequencing typically uses selection and isolation of the target gene, region, or protein by either PCR amplification (e.g., region-specific primers), hybridization-based capture, or antibody. Target enrichment can be performed at various stages of the method. For example, target RNA expression can be obtained by using target-specific primers in the reverse transcription step or by enriching a subset from a more complex library using a hybridization-based method. Examples include exome sequencing or the L1000 assay (Subramanian et al., 2017, Cell, 171; 1437-1452).Target sequencing may involve any enrichment process known to those skilled in the art. A target nucleic acid having one or both ends of a universal sequence may be referred to as a modified target nucleic acid. References to nucleic acids, such as target nucleic acids, include both single-stranded and double-stranded nucleic acids unless otherwise specified. In one embodiment, the library is enriched using an index sequence or a number of index sequences. In some embodiments, the enrichment involves one or more index sequences bound to the same library molecule and is introduced, for example, via combinatorial indexing.

[0029] As used herein, the term “universal,” when used to describe nucleotide sequences, refers to a region of sequence common to two or more nucleic acid molecules, which also have regions of sequence distinct from each other. Universal sequences present in different members of a set of molecules, such as members of a sequencing library, can be used with a collection of universal capture sequences to enable the capture of multiple different nucleic acids. Non-limiting examples of universal capture sequences include sequences identical or complementary to P5 and P7 primers. Similarly, universal sequences present in different members of a set of molecules can be used with a collection of universal primers complementary to a portion of the universal sequence, such as universal primer binding sites, to replicate (e.g., sequence) or amplify multiple different nucleic acids. The terms “A14” and “B15” may be used to refer to universal primer binding sites. The terms “A14'” (A14 prime) and “B15'” (B15 prime) refer to the complements of A14 and B15, respectively. In the methods presented herein, any suitable universal primer binding site can be used, and it will be understood that the use of A14 and B15 is merely an exemplary embodiment. In one embodiment, the universal primer binding site is used as the site on which the universal primer (e.g., a sequencing primer for lead 1 or lead 2) anneals for sequencing.

[0030] The terms "P5" and "P7" may be used to refer to a universal capture sequence or capture oligonucleotide. The terms "P5'" (P5 prime) and "P7'" (P7 prime) refer to the complements of P5 and P7, respectively. Any suitable universal capture sequence or capture nucleotide may be used in the methods presented herein, and it will be understood that the use of P5 and P7 is only in exemplary embodiments. The use of capture nucleotides such as P5 and P7 or their complements on a flow cell is known in the art, as exemplified by the disclosures in International Publications 2007 / 010251, 2006 / 064199, 2005 / 065814, 2015 / 106941, 1998 / 044151, and 2000 / 018957. For example, any suitable forward amplification primer, whether immobilized or in solution, may be useful in the methods presented herein for the amplification of complementary sequences. Similarly, any suitable reverse amplification primer, whether immobilized or in solution, may be useful in the methods presented herein for the amplification of complementary sequences. Those skilled in the art will understand the design and use of primer sequences suitable for nucleic acid capture and / or amplification as presented herein.

[0031] As used herein, the term “primer” and its derivatives generally refer to any nucleic acid that can hybridize to a sequence of interest. Typically, a primer functions as a substrate on which a nucleotide can be polymerized by polymerase or on which a nucleotide sequence can be ligated, such as an index. In some embodiments, however, a primer can be incorporated into a synthesized nucleic acid chain, providing a site on which another primer can hybridize to prime the synthesis of a new chain complementary to the synthesized nucleic acid molecule. A primer may comprise any combination of nucleotides or their analogues. A primer may be a nucleic acid that is single-stranded, double-stranded, or contains single-stranded and double-stranded regions, and may comprise ribonucleotides, deoxyribonucleotides, their analogues, or mixtures thereof. The terms “polynucleotide” and “oligonucleotide” are used interchangeably herein. It should be understood that these terms are equivalent to analogues of any of the DNA, RNA, cDNA, or antibody-oligo complexes made from nucleotide analogues, and are applicable to single-stranded (sense or antisense, etc.) and double-stranded polynucleotides. As used herein, this term also includes cDNA, which is complementary or copied DNA produced from an RNA template, for example, by the action of reverse transcriptase. This term refers only to the primary structure of a molecule. Therefore, this term includes triple-stranded, double-stranded, and single-stranded deoxyribonucleic acid ("DNA"), as well as triple-stranded, double-stranded, and single-stranded ribonucleic acid ("RNA").

[0032] As used herein, the terms “adapter” and its derivatives, e.g., “universal adapter,” generally refer to any linear oligonucleotide that can be bound to a nucleic acid molecule of the Disclosure. In some embodiments, the adapter is substantially non-complementary to the 3' or 5' end of any target sequence present in the sample. In some embodiments, preferred adapter lengths are in the range of about 10–100 nucleotides, about 12–60 nucleotides, or about 15–50 nucleotides. Generally, the adapter may contain any combination of nucleotides and / or nucleic acids. In some embodiments, the adapter may contain one or more cleavable groups at one or more positions. In other embodiments, the adapter may contain sequences that are substantially identical to or substantially complementary to at least a portion of a primer, e.g., a universal primer. In some embodiments, the adapter may contain a barcode (also referred to herein as a tag or index) to assist in downstream error correction, identification, or sequencing. The terms “adaptor” and “adapter” are used interchangeably.

[0033] As used herein, the term “each” is intended to identify individual items within a set of items, but does not necessarily refer to all items within the set unless the context explicitly indicates otherwise.

[0034] As used herein, the term “transport” refers to the movement of molecules through a fluid. This term may include passive transport, such as the movement of molecules along their concentration gradient (e.g., passive diffusion). This term may also include active transport, in which molecules can move along or against their concentration gradient. Thus, transport may involve applying energy to move one or more molecules in a desired direction or to a desired location, such as an amplification site.

[0035] As used herein, “amplification,” “to amplify,” or “amplification reaction” and their derivatives generally refer to any action or process in which at least a portion of a nucleic acid molecule is replicated or copied into at least one additional nucleic acid molecule. The additional nucleic acid molecule optionally includes a sequence that is substantially identical or substantially complementary to at least a portion of the template nucleic acid molecule. The template nucleic acid molecule may be single-stranded or double-stranded, and the additional nucleic acid molecule may be independently single-stranded or double-stranded. Amplification optionally includes linear or exponential replication of the nucleic acid molecule. In some embodiments, such amplification may be carried out using isothermal conditions, and in other embodiments, such amplification may include thermal cycling. In some embodiments, amplification is multiple amplification, including simultaneous amplification of multiple target sequences in a single amplification reaction. In some embodiments, “amplification” includes amplifying at least a portion of DNA and RNA-based nucleic acids, either individually or in combination. Amplification reactions may include any amplification processes known to those skilled in the art. In some embodiments, amplification reactions include polymerase chain reaction (PCR).

[0036] As used herein, “amplification conditions” and its derivatives generally refer to conditions suitable for amplifying one or more nucleic acid sequences. Such amplification may be linear or exponential. In some embodiments, amplification conditions may include isothermal conditions, or thermal cycling conditions, or a combination of isothermal and thermal cycling conditions. In some embodiments, polymerase chain reaction (PCR) conditions are suitable conditions for amplifying one or more nucleic acid sequences. Typically, amplification conditions refer to a reaction mixture sufficient to amplify nucleic acids such as one or more target sequences adjacent to a universal sequence, or a reaction mixture sufficient to amplify an amplified target sequence ligated to one or more adapters. Generally, amplification conditions include a catalyst for amplification, or nucleic acid synthesis, such as polymerase, primers that are somewhat complementary to the nucleic acid to be amplified, and nucleotides such as deoxyribonucleotide triphosphates (dNTPs) to facilitate primer extension when hybridized to the nucleic acid. Amplification conditions may require hybridization or annealing of primers to nucleic acids, primer extension, and denaturation steps in which the extended primers are separated from the nucleic acid sequences to be amplified. Typically, though not always, amplification conditions may include thermal cycling, but in some embodiments, amplification conditions include multiple cycles in which the annealing, extension, and separation processes are repeated. Typically, amplification conditions include Mg 2+ or Mn 2+ Examples of cations include those listed above, and may also include modifiers with varying ionic strengths.

[0037] As used herein, “re-amplification” and its derivatives generally refer to any process (referred to in some embodiments as “secondary” amplification) through which at least a portion of the amplified nucleic acid molecule is further amplified via any preferred amplification process to produce a re-amplified nucleic acid molecule. The secondary amplification does not need to be identical to the original amplification process that produced the amplified nucleic acid molecule, nor does the amplified nucleic acid molecule need to be completely identical or completely complementary to the amplified nucleic acid molecule; it is only necessary that the re-amplified nucleic acid molecule contains at least a portion of the amplified nucleic acid molecule or its complement. For example, re-amplification may involve different amplification conditions and / or the use of different primers, including target-specific primers different from those used in the primary amplification.

[0038] As used herein, the term “polymerase chain reaction” (“PCR”) refers to Mullis’s method (U.S. Patents 4,683,195 and 4,683,202) which describes a method for increasing the concentration of a segment of a target polynucleotide in a mixture of genomic DNA without cloning or purification. This process for amplifying a target polynucleotide consists of introducing a large excess of two oligonucleotide primers into a DNA mixture containing the desired target polynucleotide, followed by a series of thermal cycling steps in the presence of DNA polymerase. The two primers are complementary to the respective strands of the target double-stranded polynucleotide. First, the mixture is denatured at a higher temperature, and then the primers are annealed to their complementary sequences within the polynucleotide of the molecule of interest. After annealing, the primers are extended with polymerase to form a new pair of complementary strands. The steps of denaturation, primer annealing, and polymerase extension can be repeated many times (called thermal cycling) to obtain a highly concentrated amplified segment of the desired target polynucleotide. The length (amplicon) of the amplified segment of the desired target polynucleotide is determined by the relative positions of the primers to each other, and therefore this length is a controllable parameter. By repeating this process, the method is called PCR. The desired amplified segment of the target polynucleotide becomes the dominant nucleic acid sequence (in terms of concentration) in the mixture, and therefore these are said to be "PCR amplified." In modifications of the above method, the target nucleic acid molecule can be PCR amplified using multiple different primer pairs, and in some cases, one or more primer pairs can be used per target nucleic acid molecule, thereby forming a multiplex PCR reaction.

[0039] As defined herein, “multiplex amplification” refers to the selective and non-random amplification of two or more target sequences in a sample using at least one target-specific primer. In some embodiments, multiplex amplification is performed so that some or all of the target sequences are amplified in a single reaction vessel. A “plex” of a given multiplex amplification generally refers to the number of different target-specific sequences amplified during that single multiplex amplification. In some embodiments, the plex may be about 12 plexes, 24 plexes, 48 ​​plexes, 96 plexes, 192 plexes, 384 plexes, 768 plexes, 1536 plexes, 3072 plexes, 6144 plexes, or more. The amplified target sequences are then subjected to several different methodologies (e.g., gel electrophoresis followed by densitometry, quantification by bioanalyzer or quantitative PCR, hybridization with labeled probes, incorporation of biotinylated primers followed by detection of avidin-enzyme coupling, etc.) to the amplified target sequences. 32 Detection is also possible by incorporating a P-labeled deoxynucleotide triphosphate.

[0040] As used herein, “amplified target sequence” and its derivatives generally refer to a polynucleotide sequence prepared by amplifying a target sequence using target-specific primers and the methods provided herein. The amplified target sequence may be either of the same sense (i.e., positive) or antisense (i.e., negative) with respect to the target sequence.

[0041] As used herein, the terms “ligate,” “ligation,” and their derivatives generally refer to the process of covalently bonding two or more molecules to one another, for example, covalently bonding two or more nucleic acid molecules to one another. In some embodiments, ligation includes the bonding of nicks between adjacent nucleotides of nucleic acids. In some embodiments, ligation includes forming a covalent bond between the terminal ends of a first nucleic acid molecule and the terminal ends of a second nucleic acid molecule. In some embodiments, ligation may include forming a covalent bond between the 5' phosphate group of one nucleic acid and the 3' hydroxyl group of a second nucleic acid, thereby forming a ligated nucleic acid molecule. Generally, for the purposes of this disclosure, an amplified target sequence can be ligated to an adapter to produce an adapter-ligated amplified target sequence.

[0042] As used herein, “ligase” and its derivatives generally refer to any agent capable of catalyzing the ligation of two substrate molecules. In some embodiments, a ligase includes an enzyme capable of catalyzing the bonding of nicks between adjacent nucleotides of a nucleic acid. In some embodiments, a ligase includes an enzyme capable of catalyzing the formation of a covalent bond between the 5' phosphate of one nucleic acid molecule and the 3' hydroxyl of another nucleic acid molecule, thereby forming a ligated nucleic acid molecule. Preferred ligases include, but are not limited to, T4 DNA ligase, T4 RNA ligase, and E. coli DNA ligase.

[0043] As used herein, “ligation conditions” and its derivatives generally refer to conditions suitable for ligating two molecules to one another. In some embodiments, ligation conditions are suitable for sealing nicks or gaps between nucleic acids. As used herein, the terms “nick” or “gap” are consistent with the usage of the term in the art. Typically, nicks or gaps can be ligated in the presence of an enzyme such as a ligase at a suitable temperature and pH. In some embodiments, T4 DNA ligase can bind to nicks between nucleic acids at a temperature of about 70–72°C.

[0044] As used herein, the term “flow cell” refers to a chamber containing a solid surface through which one or more fluid reagents can flow. Examples of flow cells and associated fluid systems and detection platforms readily usable in the methods of the present disclosure are, for example, described in Bentley et al., Nature 456:53-59 (2008), International Publication No. 04 / 018497, U.S. Patent No. 7,057,026, International Publication No. 91 / 06678, International Publication No. 07 / 123744, U.S. Patent No. 7,329,492, International Publication No. 7,211,414, International Publication No. 7,315,019, International Publication No. 7,405,281, and U.S. Patent Application Publication No. 2008 / 0108082.

[0045] As used herein, the term “amplicon,” when used in relation to nucleic acids, means a product that copies a nucleic acid, having a nucleotide sequence identical to or complementary to at least a portion of the nucleotide sequence of the nucleic acid. Amplicons can be produced by any of the various amplification methods that use nucleic acids or their amplicons as templates, including, for example, polymerase elongation, polymerase chain reaction (PCR), rolling circle amplification (RCA), ligation elongation, or ligation chain reaction. An amplicon may be a nucleic acid molecule having a single copy of a particular nucleotide sequence (e.g., a PCR product) or multiple copies of a nucleotide sequence (e.g., a concatemer product of RCA). The first amplicon of the target nucleic acid is typically a complementary copy. Subsequent amplicons are copies made from the target nucleic acid or the first amplicon after the generation of the first amplicon.

[0046] As used herein, the term “amplification site” refers to a site within or on an array where one or more amplicons may be generated. An amplification site may be further configured to contain, hold, or attach at least one amplicon generated at that site.

[0047] As used herein, the term “array” refers to a group of sites that can be distinguished from one another according to their relative positions. Different molecules at different sites in an array can be distinguished from one another according to their positions within the array. Individual sites in an array may contain one or more molecules of a particular type. For example, a site may contain a single target nucleic acid molecule having a particular sequence, or a site may contain several nucleic acid molecules having the same sequence (and / or complementary sequences). Sites in an array can be different features located on the same substrate. Exemplary features include, but are not limited to, wells in the substrate, beads (or other particles) in or on the substrate, protrusions from the substrate, ridges on the substrate, or channels within the substrate. Sites in an array can be separate substrates, each having a different molecule. Different molecules attached to separate substrates can be identified according to the position of the substrate on the surface where the substrates associate, or according to the position of the substrate in a liquid or gel. Exemplary arrays in which separate substrates are arranged on a surface include, but are not limited to, those having beads in wells.

[0048] As used herein, the term “capacity,” when used in relation to sites and nucleic acid materials, means the maximum amount of nucleic acid material that can occupy a site. For example, the term may refer to the total number of nucleic acid molecules that can occupy a site under certain conditions. Other measurements may be used, for example, to include the total mass of nucleic acid material that can occupy a site under certain conditions or the total number of copies of a particular nucleotide sequence. Typically, the capacity of a site in a target nucleic acid is substantially equivalent to the capacity of a site for an amplicon of the target nucleic acid.

[0049] As used herein, the term “scavenger” refers to a material, chemical substance, molecule, or portion thereof that can attach to, retain, or bind to a target molecule (e.g., a target nucleic acid). Exemplary scavengers include, but are not limited to, capture sequences (also referred herein as capture oligonucleotides) complementary to at least a portion of the target nucleic acid, members of receptor-ligand binding pairs that can bind to the target nucleic acid (or a binding portion attached thereto) (e.g., avidin, streptavidin, biotin, lectin, carbohydrate, nucleic acid-binding protein, epitope, antibody, etc.), or chemical reagents that can form a covalent bond with the target nucleic acid (or a binding portion attached thereto).

[0050] As used herein, the term “reporter portion” may refer to any identifiable tag, label, index, barcode, or group that enables the determination of the composition, identity, and / or source of the target being investigated. In some embodiments, the reporter portion may include an antibody that specifically binds to a protein. In some embodiments, the antibody may include a detectable label. In some embodiments, the reporter may include an antibody or affinity reagent labeled with a nucleic acid tag. In one embodiment, the nucleic acid is long enough to function as a substrate for a transposomal complex. In one embodiment, the nucleic acid tag may be detectable via, for example, proximity ligation assay (PLA) or proximity extension assay (PEA), sequencing-based readout (Shahi et al. Scientific Reports volume 7, Article number: 44447, 2017), or epitope-based readout such as CITE-seq (Stoeckius et al. Nature Methods 14:865-868, 2017).

[0051] As used herein, the term “clonal population” refers to a population of nucleic acids that are homogeneous with respect to a particular nucleotide sequence. A homogeneous sequence is typically at least 10 nucleotides long, but may include longer sequences, e.g., at least 50, 100, 250, 500, or 1000 nucleotides long. A clonal population may originate from a single target nucleic acid or template nucleic acid. Typically, all nucleic acids in a clonal population have the same nucleotide sequence. It will be understood that a small number of mutations (e.g., due to amplification artifacts) may occur without deviation from clonality.

[0052] As used herein, the term “Unique Molecular Identifier” or “UMI” refers to a random, non-random, or semi-random molecular tag that may be attached to a nucleic acid. When incorporated into a nucleic acid, UMIs can be used to correct subsequent amplification bias by directly counting the Unique Molecular Identifiers (UMIs) that are sequenced after amplification.

[0053] As used herein, the term "exogenous" compound, for example, an exogenous enzyme, refers to a compound that is not typically found in a particular composition or in nature. For example, if a particular composition contains a cell lysate, the exogenous enzyme is an enzyme that is not typically found in the cell lysate or in nature.

[0054] When used herein, for example, in the context of compositions, articles, nucleic acids, or nuclei, “provide” means to produce a composition, article, nucleic acid, or nucleus, to purchase a composition, article, nucleic acid, or nucleus, or to obtain a compound, composition, article, or nucleus by other means.

[0055] The term "and / or" means one or all of the enumerated elements, or any combination of two or more of the enumerated elements.

[0056] The terms “preferred” and “preferred” refer to embodiments of the disclosure that may provide a particular benefit under certain circumstances. However, other embodiments may be preferred under the same or other circumstances. Furthermore, the description of one or more preferred embodiments does not imply that other embodiments are unhelpful, nor is it intended to exclude other embodiments from the scope of the disclosure.

[0057] The terms “comprises” and their variations shall not be limited in meaning when they appear in the description and claims.

[0058] Where described herein using terms such as “include,” “includes,” or “including,” similar embodiments described herein using the terms “consisting of” and / or “consisting essentially of” are also provided.

[0059] Unless otherwise stated, "a," "an," "the," and "at least one" are used interchangeably and mean one or more.

[0060] In this specification, an enumeration of numerical ranges by endpoints includes all numbers contained within that range (for example, 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.80, 4, 5, etc.).

[0061] In any method disclosed herein that includes separate steps, the steps may be carried out in any executable order. Furthermore, any combination of two or more steps may be carried out simultaneously.

[0062] References to “one embodiment,” “embodiment,” “a particular embodiment,” or “several embodiments” mean that certain features, configurations, compositions, or properties described in relation to this embodiment are included in at least one embodiment of this disclosure. Therefore, occurrences of such phrases in various places throughout this specification do not necessarily refer to the same embodiment of this disclosure. Furthermore, certain features, configurations, compositions, or properties may be combined in any preferred manner in one or more embodiments. [Brief explanation of the drawing]

[0063] The following detailed description of exemplary embodiments of this disclosure can be best understood in conjunction with the following drawings.

[0064] [Figure 1A] A general block diagram of a different embodiment of a general exemplary method for single-cell combinatorial indexing as disclosed herein is shown. [Figure 1B] A general block diagram of a different embodiment of a general exemplary method for single-cell combinatorial indexing as disclosed herein is shown.

[0065] [Figure 2] A schematic diagram of a single-cell combinatorial indexing method, as commonly shown in Figure 1A, is presented. For simplification, only one double-stranded target nucleic acid is shown.

[0066] [Figure 3] A general block diagram of one embodiment of a general exemplary method of single-cell combinatorial indexing as disclosed herein is shown.

[0067] [Figure 4] A general block diagram of one embodiment of a general exemplary method of single-cell combinatorial indexing as disclosed herein is shown.

[0068] [Figure 5]A schematic diagram of a single-cell combinatorial indexing method, as commonly shown in Figures 1, 3, or 4, is shown. For simplification, only one double-stranded target nucleic acid is shown.

[0069] [Figure 6] A general block diagram of one embodiment of a general exemplary method for metagenomic analysis by single-cell combinatorial indexing as disclosed herein is shown.

[0070] [Figure 7] A schematic diagram of one embodiment of a general exemplary method for preparing a sequencing library having a continuous index, as disclosed herein, is shown.

[0071] [Figure 8] A schematic diagram of one embodiment of a general exemplary method for coupling enrichment with target amplification as described herein is shown.

[0072] [Figure 9] A schematic diagram of sci-ATAC-seq3 is shown. Nuclei of 1.6 million cells from 59 fetal samples were tagged with bulk Tn5 transposase. The first two indexing steps were performed by sequential ligation of each end of the Tn5 transposase complex, and the third by PCR. The first indexing step was used as the sample index.

[0073] [Figure 10] The structure of the amplicon obtained from sci-ATAC-seq3 as described in Example 1 is shown.

[0074] [Figure 11] The project workflow described in Example 2 is shown below.

[0075] Schematic diagrams are not necessarily to scale. Similar numbers used in drawings refer to similar components, processes, etc. However, it should be understood that the use of numbers to refer to components in a given drawing is not intended to limit components in another drawing labeled with the same number. Furthermore, the use of different numbers to refer to components is not intended to indicate that components with different numbers cannot be the same as or similar to other numbered components. [Modes for carrying out the invention]

[0076] The methods provided herein can be used to prepare sequencing libraries from multiple single cells. Essentially, this includes single-nucleus sequencing of transposon-accessible chromatin (sci-ATAC, U.S. Patent No. 10,059,989), single-nucleus whole-genome sequencing (U.S. Patent Application Publication No. 2018 / 0023119), single-nucleus transcriptome sequencing (U.S. Provisional Patent Application No. 62 / 680,259 and Gunderson et al. (International Publication No. 2016 / 130704)), sci-HiC (Ramani et al., Nature Methods, 2017, 14:263-266), DRUG-seq (Ye et al., Nature Commun., 9, article number 4307), or DNA and proteins, e.g., sci-CAR (Cao et al., Science, 2018, 361(6409):1380-1385), and RNA and proteins, e.g., CITE-seq (Stoeckius et al.) Any single-nucleus or single-cell library preparation method or sequencing method can be used, including, but not limited to, any combination of analyses from (al., 2017, Nature Methods. 14(9):865-868). In one embodiment, cell atlas experiments may be performed using readouts limited to chromatin-accessible DNA, whole-cell transcriptome, a very informative limited number of mRNAs, or a combination thereof.

[0077] Provision of isolated nuclei or cells

[0078] In one embodiment, the method provided herein may include providing cells or nuclei isolated from a plurality of cells (e.g., Figure 1A, block 10, Figure 3, block 30, Figure 4, block 40, Figure 6, block 600). The cells may be from any organism and may be from any cell type or any tissue of an organism. In one embodiment, the cells may be from a tissue or a biopsy, such as a fluid biopsy. In one embodiment, the cells may be embryonic cells, for example, cells obtained from an embryo. In one embodiment, the cells or nuclei may be from cancer or diseased tissue. In one embodiment, the cells or nuclei may be immune cells, such as T cells or B cells. In one embodiment, the cells may be various different cell types obtained from a single organism. In one embodiment, various different cell types obtained from a single organism may include microbial cells, such as prokaryotic cells and / or eukaryotic cells. In one embodiment, cells from different sources, e.g., different organisms and / or different tissues, are not combined at this stage. In one embodiment, cells from different sources, e.g., different organisms and / or different tissues, are combined at this stage.

[0079] In one embodiment, multiple cells may be subsets of a larger cell population. The subset may be separated from other cells based, for example, on differences in the size, morphology, or presence or absence of identifiable molecules such as proteins or glycans on the cell surface. Methods for sorting cells are known in the art and include fluorescence-activated cell sorting, magnetoactivated cell sorting, and microfluidic cell sorting.

[0080] This method may further include dissociating cells and / or isolating the nucleus. In one embodiment, conditions are used to maintain the chromatin present in the nucleus. In one embodiment, nucleosomes present in the nucleus are depleted. Methods for depleting nucleosomes are known to those skilled in the art (U.S. Patent Application Publication 2018 / 002311).

[0081] Many different single-cell library preparation methods are known in the art (Hwang et al. Experimental & Molecular Medicine, vol. 50, Article number: 96 (2018)), including, but not limited to, the Drop-seq method, the Seq-well method, and the single-cell combinatorial indexing ("sci-") method. Companies providing single-cell products and related technologies include 10X Genomics, Takara biosciences, BD biosciences, Biorad, 1cellbio, IsoPlexis, CellSee, NanoCellect, and Dolomite. Bio is one example, but is not limited to, SCI-seq. SCI-seq is a methodological framework for uniquely labeling the nucleic acid contents of a large number of single cells or single nuclei using split pool barcoding. Typically, the number of nuclei or cells may be at least two. The upper limit depends on the actual limitations of the equipment used in other steps of the method herein (e.g., multiwell plates, number of indices). The number of nuclei or cells that may be used is not intended to be limiting and may reach billions. For example, in one embodiment, nuclei or cells may be... The number of cells may be 1,000,000,000 or less, 100,000,000 or less, 10,000,000 or less, 1,000,000 or less, 100,000 or less, 10,000 or less, 1,000 or less, 500 or less, or 50 or less. In one embodiment, the number of nuclei or cells may be at least 50, at least 500, at least 1,000, at least 10,000, at least 100,000, at least 1,000,000, at least 10,000,000, at least 100,000,000, or at least 1,000,000,000.

[0082] In these embodiments using isolated nuclei, the nuclei can be obtained by extraction and fixation. Optionally, and preferably, the method for obtaining isolated nuclei does not involve enzymatic treatment.

[0083] In one embodiment, nuclei are isolated from individual cells that are adhesive or in suspension. Methods for isolating nuclei from individual cells are known to those skilled in the art. Nuclei are typically isolated from cells present in tissue. Methods for obtaining isolated nuclei typically include preparing the tissue, isolating the nuclei from the prepared tissue, and then fixing the nuclei. In one embodiment, all steps are performed on ice.

[0084] In one embodiment, tissue preparation involves rapidly freezing the tissue in liquid nitrogen and then reducing the size of the tissue to pieces with a diameter of 1 mm or less. The size of the tissue can be reduced by applying either mincing or blunt force. Mincing can be achieved with a blade for cutting the tissue into small pieces. Applying blunt force can be achieved by crushing the tissue with a hammer or similar object, and the resulting composition of the crushed tissue is called a powder.

[0085] Nuclear isolation can be achieved by incubating a piece or powder in cell lysis buffer for at least 1 to 20 minutes, such as 5, 10, or 15 minutes. A useful buffer is one that promotes cell lysis while preserving nuclear integrity. Examples of cell lysis buffers include 10 mM Tris-HCl, pH 7.4, 10 mM NaCl, 3 mM MgCl2, 0.1% IGEPAL CA-630, 1% SUPERase In RNAse inhibitor (20 U / μL, Ambion), and 1% BSA (20 mg / mL, NEB). Standard nuclear isolation methods often use one or more exogenous compounds, such as exogenous enzymes, to aid in isolation. Examples of useful enzymes that may be present in cell lysis buffer include, but are not limited to, protease inhibitors, lysozyme, proteinase K, surfactants, lysostafin, zymolyase, cellulose, protease, or glycanase (Islam et al. Micromachines (Basel), 2017, 8(3):83; www.sigmaaldrich.com / life-science / biochemicals / biochemical-products.html?TablePage=14573107). In one embodiment, one or more exogenous enzymes are not present in the cell lysis buffer useful in the method described herein. For example, the exogenous enzymes are (i) not added to the cells before mixing with the lysis buffer, (ii) not present in the cell lysis buffer before mixing with the cells, (iii) not added to the mixture of cells and cell lysis buffer, or a combination thereof. Those skilled in the art will recognize that the concentrations of these components can be changed to some extent without reducing the usefulness of the cell lysis buffer for isolating nuclei. Next, the extracted nuclei are purified by washing with a nuclear buffer one or more times. Examples of nuclear buffers include 10 mM Tris-HCl, pH 7.4, 10 mM NaCl, 3 mM MgCl2, 1% SUPERase In RNAse inhibitor (20 U / μL, Ambion), and 1% BSA (20 mg / mL, NEB).Similar to cell lysis buffers, exogenous enzymes do not need to be present in the nuclear buffer used in the methods of this disclosure. Those skilled in the art will recognize that the concentrations of these components can be modified to some extent without reducing the usefulness of the nuclear buffer for isolating nuclei. Those skilled in the art will recognize that BSA and / or surfactants may be useful in buffers used for isolating nuclei.

[0086] The isolated nuclei can be fixed by exposure to a crosslinking agent. Useful examples of crosslinking agents include, but are not limited to, paraformaldehyde and formaldehyde. Paraformaldehyde may be present in concentrations of 1% to 8%, such as 4%. Formaldehyde may be present in concentrations of 30% to 45%, such as 37%. Treatment of nuclei with a crosslinking agent may involve adding the crosslinking agent to a suspension of nuclei and incubating at 0°C. Other methods of fixation include, but are not limited to, methanol fixation. Optionally and preferably, washing in a nucleus buffer is performed after fixation.

[0087] The isolated fixed nuclei can be immediately divided and rapidly frozen in liquid nitrogen for later use. When prepared for use after freezing, the thawed nuclei can be permeated, for example, on ice with 0.2% Triton X-100 for 3 minutes and briefly sonicated to reduce nucleus aggregation.

[0088] Conventional tissue nucleus extraction techniques typically involve incubating the tissue at a high temperature (e.g., 37°C) for 30 minutes to several hours with a tissue-specific enzyme (e.g., trypsin), followed by lysing the cells in cell lysis buffer. The nuclear isolation method described herein offers several advantages: (1) No artificial enzymes are introduced, and the entire process is performed on ice. This reduces the potential perturbation to the cellular state (e.g., chromatin tissue state or transcriptome state). (2) This novel method has been validated across most tissue types, including disease samples such as brain, lung, kidney, spleen, heart, cerebellum, and tumor tissue. Compared to conventional tissue nucleus extraction techniques that use different enzymes for different tissue types, the novel technique can potentially reduce bias when comparing cellular states from different tissues. (3) This novel method also reduces costs and increases efficiency by eliminating the enzyme processing step. (4) Compared to other nucleation techniques (e.g., Dounce tissue grinders), this new technique is more robust to different tissue types (e.g., the Dounce method requires optimizing the Dounce cycle for different tissues) and can process large specimens at high throughput (e.g., the Dounce method is limited by the size of the grinder).

[0089] Optionally, isolated nuclei may not contain nucleosomes, or they may be subjected to conditions that deplete nucleosomes and generate nucleosome-depleted nuclei.

[0090] Insertion of universal array

[0091] The methods provided herein involve inserting one or more universal sequences into nucleic acids present in the nucleus or cells. In one embodiment, the incorporation of one or more universal sequences occurs before the distribution of subsets (Figure 1A, block 11; Figure 1B, block 110), and in other embodiments, the incorporation of one or more universal sequences occurs after the distribution of subsets (Figure 3, block 32; Figure 4, block 42; block 45). In some embodiments, the index may also be combined with the universal sequences or associated with cells or nuclei as an optional step separate from the insertion of one or more universal sequences. Optional indexing of nuclei or cells may occur before or after the insertion of universal sequences (Figure 1A, block 12). In one embodiment, the samples are indexed before the distribution of a subset of nuclei or cells (Figure 1A, block 13). In some embodiments, multiple samples are indexed before the distribution of a subset of nuclei or cells (Figure 1A, block 13).

[0092] In one embodiment, a transpososome complex is used. The transpososome complex binds to a transposase recognition site and can insert the transposase recognition site into a target nucleic acid in the nucleus in a process sometimes called “tagging.” In some such insertion events, one strand of the transposase recognition site can be transferred to the target nucleic acid. Such a strand is referred to as the “transfer strand.” In one embodiment, the transpososome complex comprises a dimeric transposase having two subunits and two discontinuous transposon sequences. In another embodiment, the transposase comprises a dimeric transposase having two subunits and discontinuous transposon sequences. In one embodiment, the 5' ends of one or both strands of the transposase recognition site can be phosphorylated.

[0093] Some embodiments may include the use of highly active Tn5 transposases and Tn5-type transposase recognition sites (Goryshin and Reznikoff, J. Biol. Chem., 273:7367 (1998)), or MuA transposases and Mu transposase recognition sites containing R1 and R2 terminal sequences (Mizuuchi, K., Cell, 35:785, 1983; Savilahti, H et al., EMBO J., 14:4893, 1995). Tn5 mosaic terminal (ME) sequences can also be used by those skilled in the art.

[0094] Further examples of transposition systems that can be used with specific embodiments of the compositions and methods provided herein include Staphylococcus aureus Tn552 (Colegio et al., J. Bacteriol., 183:2384-8, 2001; Kirby C et al., Mol. Microbiol., 43:173-86, 2002), Ty1 (Devine and Boeke, Nucleic Acids Res., 22:3765-72, 1994, and International Publication No. 95 / 23875), transposon Tn7 (Craig, NL, Science. 271:1512, 1996; Craig, NL, Review in Curr Top Microbiol Immunol., 204:27-48, 1996), Tn / O and IS10 (Kleckner N et al., Curr Top Microbiol Immunol., 204:49-82, 1996), Mariner transposase (Lampe DJ, et al., EMBO J, 15:5470-9, 1996), Tc1 (Plasterk RH, Curr. Topics Microbiol. Immunol., 204:125-43, 1996), P element (Gloor, GB, Methods Mol. BiBiol, 260:97-114, 2004), Tn3 (Ichikawa and Ohtsubo, J Biol. Chem. 265:18829-32, 1990), bacterial insertion sequence (Ohtsubo and Sekine, Curr. Topics Microbiol. Immunol. 204:1-26, 1996), retrovirus (Brown et al., Proc Natl Acad Sci Examples include (USA, 86:2525-9, 1989) and yeast retrotransposons (Boeke and Corces, Annu Rev Microbiol. 43:403-34, 1989). Other examples include IS5, Tn10, Tn903, IS911, and modified versions of transposase family enzymes (Zhang et al., (2009) PLoS Genet. 5:e1000689. Epub October 16, 2009; Wilson C. et al. (2007) J. Microbiol. Methods 71:332-5).

[0095] Other examples of integrases that may be used with the methods and compositions provided herein include retroviral integrases and integrase recognition sequences of such retroviral integrases, e.g., integrases from HIV-1, HIV-2, SIV, PFV-1, and RSV.

[0096] Transposon sequences useful in the methods and compositions described herein are described in U.S. Patent Application Publication No. 2012 / 0208705, U.S. Patent Application Publication No. 2012 / 0208724, and International Publication No. 2012 / 061832. In some embodiments, the transposon sequence includes a first transposase recognition site and a second transposase recognition site.

[0097] Some transposomal complexes useful herein include transposases having two transposon sequences. In some such embodiments, the two transposon sequences are not linked to each other; in other words, the transposon sequences are not contiguous. Examples of such transpososomes are known in the art (see, for example, U.S. Patent Application Publication 2010 / 0120098).

[0098] In one embodiment, tagging is used to produce a target nucleic acid containing different universal sequences at each end (e.g., a universal primer binding site such as A14 at one end and a universal primer binding site such as B15 at the other end). This can be achieved by using two types of transposome complexes, each containing a different nucleotide sequence that is part of the transport chain. The universal sequence can serve multiple purposes. For example, but not limited to, the universal sequence can act as a complementary sequence for hybridization in a subsequent amplification step to add another nucleotide sequence (e.g., an index), can act as a site for annealing for sequencing by a universal primer (e.g., a sequencing primer for read 1 or read 2), or can act as a "landing pad" in a subsequent step to anneal a nucleotide sequence that can be used as a primer to add another nucleotide sequence, such as an index, to the target nucleic acid.

[0099] In some embodiments, the transposomal complex includes a transposon sequence nucleic acid that combines two transposase subunits to form a “loop-shaped complex” or “loop-shaped transposome.” In one embodiment, the transposome includes a dimeric transposase and a transposon sequence. The loop-shaped complex can ensure that the transposon is inserted into the target DNA while maintaining the original order information of the target DNA without fragmenting the target DNA. As understood, the loop-shaped structure may insert a desired nucleic acid sequence, such as a universal sequence, into the target nucleic acid while maintaining the physical connectivity of the target nucleic acid. In some embodiments, the transposon sequence of the loop-shaped transposomal complex may include fragmentation sites so that the transposon sequence can be fragmented to create a transposomal complex containing two transposon sequences. Such a transposomal complex is useful for ensuring that the nearby target DNA fragment into which the transposon is inserted receives a barcode combination that can be clearly assembled at a later stage of the assay. In one embodiment, an index combination is added after the insertion of one or more universal sequences into the target nucleic acid.

[0100] In one embodiment, nucleic acid fragmentation is achieved by using fragmentation sites present in the nucleic acid. Typically, fragmentation sites are introduced into the target nucleic acid by using a transposome complex. In one embodiment, after fragmentation of the nucleic acid fragment, the transposase remains bound to the nucleic acid fragment so that the nucleic acid fragments derived from the same genomic DNA molecule remain physically linked (Adey et al., 2014, Genome Res., 24:2041-2049, Amini S. et al. (2014) Nat Genet 46:1343-1349). For example, a loop-shaped transposome complex may contain fragmentation sites. Fragmentation sites can be used to cleave physical associations but cannot be used to cleave informational associations between index sequences incorporated into the target nucleic acid. Cleavage may be performed by biochemical, chemical, or other means. In some embodiments, fragmentation sites may include nucleotides or nucleotide sequences that can be fragmented by various means. Examples of fragmentation sites include, but are not limited to, restriction endonuclease sites, at least one ribonucleotide cleavable by RNAse, nucleotide analogs cleavable in the presence of specific chemical agents, diol bonds cleavable by treatment with periodate, disulfide groups cleavable with chemical reducing agents, cleavable moieties subject to photochemical cleavage, and peptides cleavable by peptidase enzymes or other suitable means (see, for example, U.S. Patent Application Publication 2012 / 0208705, U.S. Patent Application Publication 2012 / 0208724, and International Publication 2012 / 061832). In one embodiment, the transposase remains bound to the nucleic acid fragment and maintains the physical binding between nucleic acid fragments derived from the same genomic DNA molecule until removal by the use of appropriate conditions, such as the addition of a protein denaturant (e.g., SDS) or a chelating agent (e.g., EDTA). This type of approach enables the derivation of continuity information by capturing sequentially linked and transposed target nucleic acids (U.S. Patent Application Publication 2019 / 0040382). The continuity information can be preserved by maintaining the association of adjacent template nucleic acid fragments within the target nucleic acid using a transposase.

[0101] Instead of rearrangement, target nucleic acids can be obtained by fragmentation. Fragmentation of primary nucleic acids from a sample can be achieved in any order by enzymatic, chemical, or mechanical methods, after which adapters are attached to the ends of the fragments. Examples of enzymatic fragmentation include CRISPR and Talen-like enzymes, as well as enzymes that unwind DNA (e.g., helicases) that can create single-stranded regions into which the DNA fragments can hybridize and initiate extension or amplification. For example, helicase-based amplification can be used (Vincent et al., 2004, EMBO Rep., 5(8):795-800). In one embodiment, extension or amplification is initiated using random primers. Examples of mechanical fragmentation include atomization or sonication.

[0102] Mechanical fragmentation of primary nucleic acids results in fragments having a heterogeneous mixture of blunt ends, 3' overhang ends, and 5' overhang ends. Therefore, it is desirable to repair the fragment ends using methods known in the art, for example, to generate ends best suited for attaching adapters to the blunt regions. In certain embodiments, the fragment ends of a nucleic acid population are blunt ends. More specifically, the fragment ends are blunt and phosphorylated. The phosphate moiety can be introduced by enzymatic treatment, for example, using polynucleotide kinase.

[0103] In one embodiment, fragmented nucleic acids are prepared using overhang nucleotides. For example, a single overhang nucleotide can be added by a specific type of enzyme, such as Taq polymerase or Klenow exo-minus polymerase, which has template-independent terminal transferase activity that adds a single deoxyribonucleotide, such as adding nucleotide "A" to the 3' end of a DNA molecule. Using such an enzyme, a single nucleotide "A" can be added to the 3' end of the blunt end of each strand of a double-stranded nucleic acid fragment. Thus, by reaction with Taq or Klenow exo-minus polymerase, "A" can be added to the 3' end of each end-repaired strand of a double-stranded target fragment, while the adapter may be a T construct having compatible "T" overhangs present at the 3' end of each region of the double-stranded nucleic acid of the universal adapter. In one embodiment, multiple "T" nucleotides (Swift Biosciences, Ann Arbor, MI) can be added using terminal deoxynucleotidyltransferase (TdT). This type of terminal modification also prevents self-ligation of both the vector and the target, so that there is a bias towards forming target nucleic acids with the same adapter at each end.

[0104] The primary nucleic acid can be DNA, RNA, or a DNA / RNA hybrid. In embodiments where the primary nucleic acid is RNA, incorporating one or more universal sequences into the nucleic acid present in the nucleus or cell typically involves converting RNA to DNA. Various methods can be used, but in some embodiments, conventional methods used to produce cDNA are included. For example, a primer having a poly-T sequence at the 3' end and an adapter upstream of the poly-T sequence can be annealed to an mRNA molecule and extended using reverse transcriptase. This results in a one-step conversion of mRNA to DNA, and optionally, a one-step conversion of the universal sequence to the 3' end. In one embodiment, the primer may also contain one or more index sequences. In one embodiment, random primers are used.

[0105] Non-coding RNA can also be converted to DNA and optionally modified to include a universal sequence using various methods. For example, an adapter can be added using a first primer containing a random sequence and a template switch primer, and either primer can contain a universal sequence adapter. A reverse transcriptase with terminal transferase activity can be used to result in the addition of a non-template nucleotide to the 3' end of the synthetic strand, and the template switch primer contains a nucleotide to anneal with the non-template nucleotide added by the reverse transcriptase. An example of a useful reverse transcriptase is Moroni mouse leukemia virus reverse transcriptase. In certain embodiments, the SMARTer® reagent (catalog no. 634926), available from Takara Bio USA, Inc., is used to add a universal sequence to non-coding RNA and, optionally, to mRNA for use as a template switch. Optionally, the template switch primer can be used with RNA in combination with a primer having a poly-T sequence to add a universal sequence to both ends of the DNA target nucleic acid produced from the RNA.

[0106] Subset distribution

[0107] The methods provided herein involve distributing a subset of isolated nuclei or cells into multiple compartments (Figure 1A, Block 13; Figure 1B, Block 115; Figure 3, Block 31; Figure 4, Block 41; Block 44). The methods may include multiple distribution steps for dividing a population of isolated nuclei or cells (also referred to herein as a pool) into subsets. Typically, a subset of isolated nuclei or cells, e.g., a subset residing in multiple compartments, is indexed with a compartment-specific index and then pooled. Thus, the methods typically include at least one “split and pool” step, obtaining pooled isolated nuclei or cells, distributing them, and adding compartment-specific indices; the number of “split and pool” steps may depend on the number of different indices added to the target nucleic acids. Each initial subset of nuclei or cells before indexing may be unique, distinct from other subsets. For example, each first subset may come from a unique sample, such as a unique organism or unique tissue. After indexing, subsets can be pooled, split into subsets, and pooled again as needed until a sufficient number of indices have been added to the target nucleic acids. This process assigns a unique index or combination of indices to each single cell or single nucleus, resulting in combinatorial indexing as described herein. After indexing is complete, for example, after the addition of one, two, three, or more indices, the isolated nuclei or cells can be lysed. In some embodiments, indexing and lysis may occur simultaneously.

[0108] The number of nuclei or cells in a subset, and therefore in each compartment, can be at least 1. In one embodiment, the number of nuclei or cells in a subset is 100,000,000 or less, 10,000,000 or less, 100,000 or less, 10,000 or less, 4,000 or less, 3,000 or less, 2,000 or less, 1,000 or less, 500 or less, or 50 or less. In one embodiment, the number of nuclei or cells in a subset can be 1 to 1,000, 1,000 to 10,000, 10,000 to 100,000, 100,000 to 1,000,000, 1,000,000 to 10,000,000, or 10,000,000 to 100,000,000. In one embodiment, the number of nuclei or cells present in each subset is approximately equal. The number of nuclei or cells present in each subset, and therefore in each compartment, is based in part on the desire to reduce index collisions, which are the presence of two nuclei or cells having the same index combination ending in the same compartment in this step of the method. Methods for distributing nuclei or cells into subsets are known and routine to those skilled in the art. Fluorescence-activated cell sorting (FACS) cytometry can be used, but in some embodiments, the use of simple dilutions is preferred. In one embodiment, FACS cytometry is not used. Optionally, nuclei of different ploidy levels can be gated and enriched by staining, e.g., DAPI (4',6-diamidino-2-phenylindole) staining. Staining can also be used to identify single cells from doublets during sorting.

[0109] The number of compartments in the distribution process (and subsequent indexing) may depend on the format used. For example, the number of compartments may be 2 to 96 (when using a 96-well plate), 2 to 384 (when using a 384-well plate), or 2 to 1536 (when using a 1536-well plate). In one embodiment, multiple plates may be used. Examples of compartments include, but are not limited to, wells, droplets, and microfluidic compartments. In one embodiment, each compartment may be a droplet. If the type of compartment used is a droplet containing two or more nuclei or cells, any number of droplets may be used, such as at least 10,000, at least 100,000, at least 1,000,000, or at least 10,000,000 droplets. The isolated subset of nuclei or cells is typically indexed within the compartment before pooling.

[0110] Combinatorial Indexing

[0111] The methods provided herein include adding compartment-specific indices to nuclei or cells present in a sample (Figure 1B, block 112), or adding compartment-specific indices to isolated subsets of nuclei or cells distributed to different compartments (e.g., Figure 1A, block 14, Figure 3, block 32, Figure 4, blocks 42 and 45, Figure 6, block 601). In some embodiments, universal sequences may also be incorporated along with the indices. Index sequences, also called tags or barcodes, are useful as characteristic markers for compartments where specific nucleic acids are present. Thus, in some embodiments, the index is a nucleic acid sequence tag bound to each of the target nucleic acids present in a particular compartment, and its presence is used to indicate or identify the compartment where a population of nuclei or cells resides at a particular stage of the method.

[0112] In one embodiment, multiple indices are added. The incorporation of each index occurs in a single split and pool indexing. One, two, three, or more split and pool barcodings result in single, double, triple, or multiple (e.g., quadruple or more) indexed target nucleic acids.

[0113] Indexes may be added to one or both ends of a target nucleic acid. For example, a modified target nucleic acid having two or more indexes may have different indexes at each end (an example is shown in Figure 5A). In Figure 5A, the target nucleic acid 55 is modified to have four distinct indexes, two at one end (51 and 52) and two at the other end (53 and 54). In other embodiments, the modified target nucleic acid may have grouped indexes at one or both ends (an example is shown in Figure 5B). In Figure 5B, the target nucleic acid 56 is modified to have four distinct indexes at each end (51, 52, 53, and 54). A set of indexes present at one end of a target nucleic acid may be referred to as a “continuous index”. In one embodiment, a continuous index has no nucleotides between each index. In other embodiments, there may be one, two, three, four, or more nucleotides between one or more indices in a continuous index. As described herein, continuous indexes may be useful in identifying members of a library having a particular set of indexes. For example, a continuous index can promote the enrichment of library members originating from the same cell.

[0114] The index sequence can be any suitable number of nucleotide lengths, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides. Four nucleotide tags allow for the multiplexing of 256 samples in the same array, while six nucleotide tags enable the processing of 4096 samples in the same array.

[0115] In one embodiment, the index is added after the universal sequence has been incorporated into the nuclear or cellular DNA nucleic acid, for example, by a transposomal complex. The incorporation of the index sequence may essentially involve a process comprising one, two, or more steps, using any combination of ligation, extension, hybridization, adsorption, specific or nonspecific primer interactions, or amplification. In one embodiment, the index is added during cDNA synthesis. In one embodiment, the index is added through tagging. The nucleotide sequences added to one or both ends of the target nucleic acid may also include one or more universal sequences and / or other useful sequences, such as unique molecular identifiers.

[0116] Various methods can be used to add an index to a nucleic acid containing a universal sequence, and the method of index addition is not intended to be limited. In one embodiment, the target nucleic acid has different universal sequences at each end (e.g., A14 at one end and B15 at the other), and those skilled in the art will recognize that a specific sequence can be added to one or both ends of the target nucleic acid. The universal sequence added by the transposome complex can be used as a “landing pad” in a subsequent step of annealing a nucleotide sequence, which can be used as a primer for adding another nucleotide sequence to the target nucleic acid, such as another index and / or another universal sequence. For example, in one embodiment, the incorporation of an index sequence involves ligating primers to one or both ends of the nucleic acid. Ligation of the primer can be assisted by the presence of the universal sequence at each end of the target nucleic acid. An example of a primer is a double hairpin ligation. A double ligation can be ligated to one or preferably both ends of the target nucleic acid.

[0117] In one embodiment, blunt-end ligation can be used. In another embodiment, the target nucleic acid is prepared using a single overhang nucleotide by the activity of a specific type of DNA polymerase, such as Taq polymerase or Klenow exo-minus polymerase having template-independent end-transferase activity, which adds one or more deoxynucleotides, such as deoxyadenosine (A), to the 3' end of the target nucleic acid. In some cases, the overhang nucleotide is two or more bases. Using such an enzyme, a single nucleotide "A" can be added to the blunt end 3' end of each strand of the target nucleic acid. Thus, by reaction with Taq or Klenow exo-minus polymerase, "A" can be added to the 3' end of each strand of a double-stranded target fragment, while any further sequence added to each end of the target nucleic acid may include a compatible "T" overhang present at the 3' end of each region of the double-stranded nucleic acid to be modified. This end modification also prevents self-ligation of the nucleic acid, such that there is a bias to form an indexed target nucleic acid adjacent to the sequence added in this embodiment.

[0118] In one embodiment, index incorporation is performed by an exponential amplification reaction such as PCR. The universal sequence located at the end of the target nucleic acid acts as a primer and can be used for annealing sequences that can be extended in the amplification reaction.

[0119] The index and other useful sequences can be added in a single step or in multiple steps. For example, the index and any other useful sequences can be added by ligation or extension, or a two-step method can be used, for example, which includes ligating the universal sequence and then amplifying the universal sequence to further modify it to include the index and any other useful sequences.

[0120] In one embodiment, a universal sequence useful for immobilizing and / or sequencing the target nucleic acid is added by the addition of a sequence during the indexing process. In another embodiment, the indexed target nucleic acid can be further processed to add a universal sequence useful for immobilizing and sequencing the target nucleic acid. Those skilled in the art will recognize that in embodiments where the compartments are droplets, the sequence for immobilizing the nucleic acid fragment is optional. In one embodiment, the incorporation of a universal sequence useful for immobilizing and sequencing the fragment involves ligating the 5' and 3' ends of the indexed nucleic acid fragment with the same universal adapter (also called a “mismatch adapter,” whose general features are described in U.S. Patent No. 7,741,463 (Gormley et al.) and U.S. Patent No. 8,053,192 (Bignell et al.)). In one embodiment, the universal adapter includes all sequences necessary for sequencing, including a sequence for immobilizing the indexed nucleic acid fragment on the array.

[0121] The resulting indexed fragments collectively provide a library of nucleic acids that can be immobilized and then sequenced. The term "library," also referred to herein as a sequencing library, refers to an assembly of nucleic acid fragments from a single nucleus or single cell containing various combinations of known universal sequences and indices at the 3' and 5' ends. A library may include, for example, nucleic acids from accessible DNA, a whole genome, or a whole transcriptome, nucleic acids representing a specific protein, or combinations thereof, and can be used for sequencing.

[0122] Indexed nucleic acid fragments can be subjected to selection conditions for a predetermined size range, such as 150–400 nucleotides in length, such as 150–300 nucleotides. The resulting indexed nucleic acid fragments can be pooled and subjected to a cleanup process to improve the purity of the DNA molecules by optionally removing at least some of the unbound universal adapters or primers. Any suitable cleanup process, such as electrophoresis or size exclusion chromatography, may be used. In some embodiments, solid-phase reversible paramagnetic beads may be used to separate the desired DNA molecules from the unbound universal adapters or primers and select the nucleic acids based on size. Solid-phase reversible paramagnetic beads are commercially available from Beckman Coulter (Agencourt AMPure XP), Thermo Fisher (MagJet), Omega Biotech (Mag-Bind), Promega Beads (Promega), and Kapa Biosystems (Kapa Pure Beads).

[0123] A non-limiting exemplary embodiment of the present disclosure is shown in Figure 1A. In this embodiment, the method comprises providing a plurality of nuclei or cells (Figure 1A, block 10). The plurality of nuclei or cells may be from a sample or a plurality of samples. The method further comprises incorporating one or more universal sequences into nucleic acids present in the nuclei or cells (Figure 1A, block 11). Optionally, the method may also comprise associating an index with the nuclei or cells (e.g., nucleus or cell hashing, see International Publication No. 2020 / 180778), and in one embodiment, an index can be added to the nucleic acid by association (Figure 1A, block 12). In one embodiment, two different universal sequences are added to ultimately obtain a target nucleic acid having different universal sequences at each end. The method further comprises distributing a subset of nuclei or cells, incorporating a universal sequence into the nucleic acids located therein, and optionally incorporating at least one index into a plurality of compartments (Figure 1, block 13). The nucleic acids present in each compartment are indexed (Figure 1A, block 14), and then the nuclei or cells are pooled (Figure 1A, block 15). After the addition of a single index, the library of nucleic acids in the nuclei or cells can be further processed to prepare for sequencing (Figure 1A, block 16). However, in some preferred embodiments, it is desirable to add a second, third, or more indexes. In one embodiment, the addition of each index may include a “split and pool” step in which indexing occurs after splitting, for example, distributing a subset of nuclei or cells into multiple compartments (Figure 1A, block 13), indexing the nucleic acids present in each compartment (Figure 1A, block 14), and then pooling the nuclei or cells (Figure 1A, block 15). The “split and pool” step may result in indexing only one or both ends of the nucleic acids present in the nuclei or cells. After the final indexing, the library of nuclear or intracellular nucleic acids can be pooled and further processed to prepare it for sequencing (which may be inclusive sequencing or targeted sequencing) (Figure 1A, Block 16).

[0124] Another non-limiting exemplary embodiment of the present disclosure is shown in Figure 1B. In this embodiment, the method includes first providing a plurality of samples to be processed in parallel (Figure 1B, block 110). The method includes incorporating one or more universal sequences into nucleic acids present in the nucleus or cells (Figure 1B, block 111), followed by indexing the nucleic acids (Figure 1B, block 112), the indexes added to each sample being unique and can be used as sample indices to identify nucleic acids derived from a particular sample. In one embodiment, two different universal sequences are added to ultimately obtain a target nucleic acid having different universal sequences at each end. The method further includes pooling nuclei or cells (Figure 1B, block 113). In one embodiment, after the addition of one index, the library of nucleic acids in the nucleus or cells can be further processed to prepare for sequencing (Figure 1B, block 114). However, in some preferred embodiments, it is desirable to add a second, third, or more indexes. In one embodiment, the addition of each index may include a “split and pool” step in which indexing occurs after splitting, for example, distributing a subset of nuclei or cells into multiple compartments (Figure 1B, block 115), indexing the nucleic acids present in each compartment (Figure 1B, block 116), and then pooling the nuclei or cells (Figure 1B, block 117). The “split and pool” step may result in indexing only one or both ends of the nucleic acids present in the nuclei or cells. After the final indexing, the library of nucleic acids from the nuclei or cells can be pooled and further processed to prepare for sequencing (which may be inclusive sequencing or targeted sequencing) (Figure 1B, block 118).

[0125] Another non-limiting exemplary embodiment of the present disclosure is shown in Figure 2. In this embodiment, the method involves using tagging to incorporate two universal sequences into a nucleic acid present in a nucleus or cell, followed by three subsequent indexing steps (Figure 2A). One transposome complex 21 contains the universal sequence 23 (e.g., A14), and another transposome complex 22 contains the universal sequence 24 (B15). Insertion of the universal sequences into the nucleic acid occurs in a bulk of multiple nuclei or cells. Figure 2A also shows the result of inserting the two universal sequences 23 and 24 into a target nucleic acid 25. The multiple nuclei or cells are distributed into different compartments, and a polynucleotide 26 containing the index is added to one side of the nucleic acid 25 by ligation using a nucleotide complementary to one of the universal sequences (e.g., A14) (Figure 2B). Multiple nuclei or cells are pooled and then distributed into different compartments, and a different polynucleotide 27 containing a second index is added to the other side of nucleic acid 25 by ligation using a nucleotide complementary to the other universal sequence (e.g., B15) (Figure 2C). Multiple nuclei or cells containing the dual-indexed nucleic acid are pooled and then distributed into different compartments, and then subjected to a PCR amplification reaction in which a polynucleotide 28 containing a third index is added to one side of nucleic acid 25, and a polynucleotide 29 containing a fourth index is added to the other side of nucleic acid 25 (Figure 2D). After the addition of the last index, the library of nucleic acids in the nuclei or cells can be pooled and further processed to prepare for sequencing (which may be inclusive sequencing or targeted sequencing).

[0126] Another non-limiting exemplary embodiment of the present disclosure is shown in Figure 3. In this embodiment, the method comprises providing a plurality of nuclei or cells (Figure 3, block 30). The method further comprises distributing a subset of nuclei or cells into a plurality of compartments (Figure 3, block 31). The nucleic acids present in the nuclei or cells of each compartment are modified by incorporating an index and / or a universal sequence (Figure 3, block 32). In another embodiment, the nucleic acids present in the nuclei or cells of each compartment are modified by incorporating the same universal sequence (e.g., tagging using transposons having the same universal sequence), followed by the addition of a compartment-specific index. The nuclei or cells are then pooled (Figure 3, block 33). After the addition of the index and / or universal sequence, the library of nucleic acids in the nuclei or cells can be further processed to prepare for sequencing (Figure 3, block 34). However, in some preferred embodiments, it is desirable to add a second, third, or more indexes. Optionally, a universal sequence may also be added. Each indexing step may involve a “split and pool” process in which indexing occurs after splitting, for example, distributing a subset of nuclei or cells into multiple compartments (Figure 3, block 31), indexing the nucleic acids present in each compartment (Figure 3, block 32), and then pooling the nuclei or cells (Figure 3, block 33). The “split and pool” step may result in indexing being applied to only one or both ends of the nucleic acids present in the nuclei or cells. After the final indexing, the library of nucleic acids from the nuclei or cells can be pooled and further processed to prepare for sequencing (which may be inclusive sequencing or targeted sequencing) (Figure 3, block 34).

[0127] A more non-limiting exemplary embodiment of the present disclosure is shown in Figure 4. In this embodiment, the method includes RNA analysis. A plurality of nuclei or cells are provided (Figure 4, block 40), which can be obtained from a sample or a plurality of samples. A subset of nuclei or cells is distributed into a plurality of compartments (Figure 4, block 41). Optionally, the method may also include associating an index with the nuclei or cells (e.g., nucleus or cell hashing, see International Publication No. 2020 / 180778) or nucleic acids before distribution. The nucleic acids present in the nuclei or cells of each compartment are modified using reverse transcriptase to insert an index and / or universal sequence (Figure 4, block 42), and then the nuclei or cells are pooled (Figure 4, block 43). The method further includes distributing a subset of nuclei or cells into a plurality of compartments (Figure 4, block 44). The nucleic acids present in the nuclei or cells of each compartment are modified by inserting another index and / or universal sequence (Figure 4, block 45), and then the nuclei or cells are pooled (Figure 4, block 46). After the addition of the index and / or universal sequence, the library of nucleic acids from the nucleus or cells can be further processed to prepare it for sequencing (Figure 4, block 47). However, in some preferred embodiments, it is desirable to add a third, fourth, or more index. Optionally, a universal sequence can also be added. Each index addition may involve a “split and pool” step in which indexing occurs after splitting, for example, distributing a subset of nuclei or cells into multiple compartments (Figure 4, block 44), indexing the nucleic acids present in each compartment (Figure 4, block 45), and then pooling the nuclei or cells (Figure 4, block 46). The “split and pool” step may result in indexing being added to only one or both ends of the nucleic acids present in the nuclei or cells. After the final index addition, the library of nucleic acids from the nucleus or cells can be pooled and further processed to prepare it for sequencing (which may be inclusive sequencing or targeted sequencing) (Figure 4, block 47).

[0128] Preparation of fixed samples for sequencing

[0129] Methods for immobilizing indexed fragments from one or more sources onto a substrate are known in the art. In one embodiment, the indexed fragment is enriched using a plurality of capture sequences having specificity for the indexed fragment, and the capture sequences can be immobilized on the surface of a solid substrate. For example, the capture sequence may include a first member of a binding pair (e.g., P5'), and a second member of the binding pair (P5) is immobilized on the surface of the solid substrate. Similarly, methods for amplifying immobilized indexed fragments include, but are not limited to, bridge amplification and binding equilibrium exclusion. Methods for immobilization and amplification before sequencing are described, for example, in Bignell et al. (U.S. Patent No. 8,053,192), Gunderson et al. (International Publication No. 2016 / 130704), Shen et al. (U.S. Patent No. 8,895,249), and Pipenburg et al. (U.S. Patent No. 9,309,502).

[0130] Pooled samples can be immobilized during preparation for sequencing. Sequencing can be performed as an array of single molecules or amplified before sequencing. Amplification can be performed using one or more immobilized primers. The immobilized primers may be, for example, on a plane or on a pool of beads. The pool of beads may be isolated in an emulsion having a single bead in each "compartment" of the emulsion. At concentrations of only one template per "compartment," only a single template is amplified on each bead.

[0131] As used herein, the term “solid-phase amplification” refers to any nucleic acid amplification reaction carried out on or in association with a solid support such that all or part of the amplified product is immobilized on the solid support during formation. Specifically, the term encompasses solid-phase polymerase chain reactions (solid-phase PCR) and solid-phase isothermal amplification, which are reactions similar to standard solution-phase amplification, except that one or both of the forward and reverse amplification primers are immobilized on the solid support. Solid-phase PCR involves systems such as colonization in emulsions where one primer is immobilized on a bead and the other is in free solution, or in solid-phase gel matrices where one primer is immobilized on the surface and the other is in free solution.

[0132] In some embodiments, the solid support includes a patterned surface. “Patterned surface” refers to an arrangement of different regions within or on the exposed layer of the solid support. For example, one or more regions may be features where one or more amplification primers are present. These features may be separated by interstitial regions where amplification primers are absent. In some embodiments, the pattern may be an xy format of features in rows and columns. In some embodiments, the pattern may be a repeating arrangement of features and / or interstitial regions. In some embodiments, the pattern may be a random arrangement of features and / or interstitial regions. Exemplary patterned surfaces that can be used in the methods and compositions described herein are described in U.S. Patents No. 8,778,848, 8,778,849, and 9,079,148, and U.S. Patent Application Publication No. 2014 / 0243224.

[0133] In some embodiments, the solid support comprises an array of wells or depressions on its surface. This can be manufactured using a variety of techniques, including but not limited to photolithography, stamping, molding, and microetching, as is commonly known in the art. As is understood in the art, the technique used depends on the composition and shape of the array substrate.

[0134] The features within the patterned surface may be wells in an array of wells (e.g., microwells or nanowells) on another suitable solid support with a patterned covalent gel, such as glass, silicon, plastic, or poly(N-(5-azidoacetamylpentyl)acrylamide-coacrylamide) (PAZAM, see, e.g., U.S. Patent Application Publication 2013 / 184796, International Publication 2016 / 066586, and 2015 / 002813). This process creates a gel pad used for sequencing, which can be stable over sequencing operations for a large number of cycles. Covalently bonding the polymer to the wells is useful for maintaining the gel in the structured feature area throughout the lifetime of the structured substrate during various applications. However, in many embodiments, the gel does not need to be covalently bonded to the wells. For example, under some conditions, silane-free acrylamide (SFA, see, e.g., U.S. Patent No. 8,563,477) that is not covalently bonded to any part of the structured substrate can be used as the gel material.

[0135] In a specific alternative embodiment, the structured substrate can be prepared by patterning a solid support material using wells (e.g., microwells or nanowells), coating the patterned support with a gel material (e.g., PAZAM, SFA, or a chemically modified variant thereof), such as an azidated form of SFA (azido-SFA), and polishing the gel-coated support, for example, by chemical polishing or mechanical polishing, thereby retaining the gel within the wells but removing or inactivating substantially all of the gel from the interstitial regions of the structured substrate surface between the wells. Primer nucleic acids can then be attached to the gel material. Next, a solution of indexed fragments can be brought into contact with the polished substrate so that individual indexed fragments are seeded into individual wells via interaction with primers attached to the gel material, but the target nucleic acids do not occupy the interstitial region because the gel material is absent or inactive. Amplification of the indexed fragments would be limited to the wells because the absence or inactivity of the gel in the interstitial region prevents outward migration of the growing nucleic acid colonies. The process is conveniently manufacturable, scalable, and utilizes conventional micro or nano fabrication methods.

[0136] This disclosure encompasses “solid-phase” amplification methods in which only one amplification primer is immobilized (other primers are typically present in free solution), but in one embodiment, it is desirable that the solid support provides both immobilized forward and reverse primers. In practice, since amplification processes require excess primers to maintain amplification, there will be “multiple” identical forward primers and / or “multiple” identical reverse primers immobilized on the solid support. References to forward and reverse primers herein should be interpreted as encompassing “multiple” such primers unless the context indicates otherwise.

[0137] As those skilled in the art will understand, any given amplification reaction requires at least one type of forward primer and at least one type of reverse primer that are specific to the template being amplified. However, in certain embodiments, the forward and reverse primers may contain template-specific portions of the same sequence and may have completely identical nucleotide sequences and structures (including any non-nucleotide modifications). In other words, solid-phase amplification can be performed using only one type of primer, and such single-primer methods are included within the scope of this disclosure. Other embodiments may use forward and reverse primers that contain the same template-specific sequence but differ in some other structural features. For example, one type of primer may contain non-nucleotide modifications that are not present in the other.

[0138] In all embodiments of this disclosure, the solid-phase amplification primer is preferably immobilized to a solid support by a single-point covalent bond at or near its 5' end, allowing the template-specific portion of the primer to freely anneal to its homologous template and 3' hydroxyl groups that do not include primer extension. Any suitable covalent bonding means known in the art can be used for this purpose. The selected adhesion chemical depends on the properties of the solid support and any derivatization or functionalization applied thereto. The primer itself may include a portion that is non-nucleotide chemically modified to promote adhesion. In certain embodiments, the primer may include a sulfur-containing nucleophile, such as a phosphorothioate or thiophosphate, at its 5' end. In the case of a solid-supported polyacrylamide hydrogel, this nucleophile binds to a bromoacetamide group present in the hydrogel. A more specific means of binding the primer and template to the solid support is via a 5' phosphorothioate bond to a hydrogel consisting of polymerized acrylamide and N-(5-bromoacetamidoylpentyl)acrylamide (BRAPA), as described in International Publication No. 05 / 065814.

[0139] Certain embodiments of this disclosure may utilize a solid support comprising an inert substrate or matrix (e.g., a glass slide, polymer beads, etc.) that has been “functionalized” by the application of a layer or coating of an intermediate material containing reactive groups that enable covalent bonding to biomolecules, such as polynucleotides. Examples of such supports include, but are not limited to, polyacrylamide hydrogels supported on an inert substrate such as glass. In such embodiments, the biomolecules (e.g., polynucleotides) may be directly covalently bonded to the intermediate material (e.g., hydrogel), or the intermediate material may be noncovalently bonded to the substrate or matrix (e.g., glass substrate) itself. The term “covalent bonding to a solid support” should be interpreted as appropriate to encompass this type of arrangement.

[0140] The pooled samples may be amplified on beads, each containing forward and reverse amplification primers. In certain embodiments, a library of indexed fragments is used to prepare clustered arrays of nucleic acid colonies by solid-phase amplification, more specifically by solid-phase isothermal amplification, as described in U.S. Patent Application Publication No. 2005 / 0100900, U.S. Patent No. 7,115,400, International Publication Nos. 00 / 18957 and 98 / 44151. The terms “cluster” and “colony” are used interchangeably herein and refer to distinct sites on a solid support containing multiple identical immobilized nucleic acid strands and multiple identical immobilized complementary nucleic acid strands. The term “clustered array” refers to an array formed from such clusters or colonies. In this context, the term “array” should not be understood as requiring an ordered arrangement of clusters.

[0141] The terms “solid phase” or “surface” are used to mean either a planar array to which primers are attached, such as a flat surface, e.g., glass, silica, or plastic microscope slide, or a similar flow cell device, or beads, to which one or two primers are attached and the beads are amplified, or an array of beads on a surface after the beads have been amplified.

[0142] The clustered sequences can be prepared using a thermal cycling process, such as that described in International Publication No. 98 / 44151, or a process in which the temperature is kept constant and stretching and denaturation cycles are performed using changes in reagents. Such isothermal amplification methods are described in International Publication No. 02 / 46456 and U.S. Patent Application Publication No. 2008 / 0009420. Due to the lower temperatures useful in the isothermal process, this is particularly preferred in some embodiments.

[0143] It will be understood that any amplification method described herein or commonly known in the art may be used with universal or target-specific primers to amplify immobilized DNA fragments. Suitable amplification methods include, but are not limited to, polymerase chain reaction (PCR), strand displacement amplification (SDA), transcription-mediated amplification (TMA), and nucleic acid sequence-based amplification (NASBA), as described in U.S. Patent No. 8,003,354. One or more target nucleic acids can be amplified using these amplification methods. For example, immobilized DNA fragments can be amplified using PCR, such as multiplex PCR, SDA, TMA, and NASBA. In some embodiments, primers specifically directed to the target polynucleotide are included in the amplification reaction.

[0144] Other suitable methods for amplifying polynucleotides may include oligonucleotide extension and ligation, rolling circle amplification (RCA) (Lizardi et al., Nat. Genet. 19:225-232 (1998)), and oligonucleotide ligation assay (OLA) techniques (see generally U.S. Patents Nos. 7,582,420, 5,185,243, 5,679,524, and 5,573,907; European Patents Nos. 0320 308(B1), 0336 731(B1), and 0439 182(B1); International Publications Nos. 90 / 01069, 89 / 12696, and 89 / 09835). It will be understood that these amplification methods may be designed to amplify immobilized DNA fragments. For example, in some embodiments, the amplification method may include a ligation probe amplification or oligonucleotide ligation assay (OLA) reaction containing a primer specifically directed to the nucleic acid of interest. In some embodiments, the amplification method may include a primer extension ligation reaction containing a primer specifically directed to the nucleic acid of interest. Non-limiting examples of primer extension and ligation primers that can be specifically designed to amplify the nucleic acid of interest include primers used in the GoldenGate assay, as exemplified by U.S. Patents No. 7,582,420 and No. 7,611,869 (Illumina, San Diego, California).

[0145] DNA nanoblocks can also be used in combination with the methods and compositions described herein. Methods for creating and using DNA nanoblocks for genome sequencing can be found, for example, in U.S. Patents and Publications No. 7,910,354, No. 2009 / 0264299, No. 2009 / 0011943, No. 2009 / 0005252, No. 2009 / 0155781, and No. 2009 / 0118488, as described, for example, in Drmanac et al., 2010, Science 327(5961):78-81. In short, after fragmentation of a genomic library DNA, an adapter is ligated to the fragment, and the adapter circulates the ligated fragment by ligation with a circle ligase to perform rolling circle amplification (as described in Lizardi et al., 1998. Nat. Genet. 19:225-232 and U.S. Patent Application Publication No. 2007 / 0099208(A1)). The extended concatemer structure of the amplicon promotes coiling, thereby generating compact DNA nanoballs. The DNA nanoballs can preferably be captured on a substrate to form an ordered or patterned sequence, such that the distance between each nanoball is maintained, thereby enabling sequencing of distinct DNA nanoballs. In some embodiments, the adapter ligation, amplification, and digestion, performed sequentially, are carried out before circulation to produce a head-tail construct having several genomic DNA fragments separated by the adapter sequence.

[0146] Examples of isothermal amplification methods that may be used in the methods of this disclosure include, but are not limited to, multiple substitution amplification (MDA) as exemplified by isothermal chain substitution nucleic acid amplification as exemplified by Dean et al., Proc.Natl.Acad.Sci.USA 99:5261-66 (2002), or, for example, U.S. Patent No. 6,214,587. Other non-PCR methods that may be used in this disclosure include, for example, chain substitution amplification (SDA) as described by Walker et al., Molecular Methods for Virus Detection, Academic Press, 1995, U.S. Patents No. 5,455,166 and 5,130,238, and hyperbranched chain substitution amplification as described by Walker et al., Nucl.Acids Res.20:1691-96 (1992), or, for example, Lage et al., Genome Res.13:294-307 (2003). Isothermal amplification can be used, for example, with strand-displacement Phi 29 polymerase or Bst DNA polymerase for large fragments, and with 5'->3' exos for random primer amplification of genomic DNA. The use of these polymerases takes advantage of their high processability and strand-displacement activity. Due to their high processability, polymerases can produce fragments of 10-20 kb in length. As mentioned above, smaller fragments can be produced under isothermal conditions using polymerases with low processability and polymerases with strand-displacement activity, such as Klenow polymerase. Further descriptions of amplification reactions, conditions, and components are detailed in the disclosure of U.S. Patent No. 7,670,810.

[0147] Another polynucleotide amplification method useful in this disclosure is tagged PCR, which uses a population of two-domain primers having a random 3' region following a 5' region, as described, for example, by Grothues et al. in Nucleic Acids Res. 21(5):1321-2 (1993). The first round of amplification is performed to enable numerous initiations on thermally denatured DNA based on individual hybridization from randomly synthesized 3' regions. Due to the nature of the 3' region, the start sites are considered to be random throughout the genome. Subsequently, unbound primers may be removed, and further replication may be performed using primers complementary to a certain 5' region.

[0148] In some embodiments, isothermal amplification can be performed using binding equilibrium exclusion amplification (KEA), also known as exclusion amplification (ExAmp). The nucleic acid libraries of this disclosure can be prepared using a method comprising reacting amplification reagents to produce a number of amplification sites, each containing a substantially clonal population of amplicons from individual target nucleic acids seeded at sites. In some embodiments, the amplification reaction proceeds until a sufficient number of amplicons are produced to fill the volume of each amplification site. Thus, filling an already seeded site to volume inhibits the target nucleic acid from landing and amplifying at that site, thereby producing a clonal population of amplicons at that site. In some embodiments, apparent clonality can be achieved even if the amplification site is not filled to volume before a second target nucleic acid reaches that site. Under some conditions, amplification of the first target nucleic acid may proceed to a point where a sufficient number of copies are produced to effectively exceed or overwhelm the production of copies from the second target nucleic acid transported to that site. For example, in an embodiment using a bridge amplification process on a circular feature region with a diameter of less than 500 nm, it was determined that after 14 cycles of exponential amplification for the first target nucleic acid, contamination from the second target nucleic acid at the same site generated an insufficient number of contaminating amplicons to adversely affect sequence synthesis analysis on the Illumina sequencing platform.

[0149] In some embodiments, the amplification sites in the array may, but do not necessarily, be completely clones. Rather, in some applications, individual amplification sites may be primarily composed of amplicons from a first indexed fragment and also have low levels of contaminating amplicons from a second target nucleic acid. An array may have one or more amplification sites with low levels of contaminating amplicons, as long as the level of contamination does not have an unacceptable effect on the subsequent use of the array. For example, if the array is used in a detection application, an acceptable level of contamination is one that does not affect the signal-to-noise ratio or resolution of the detection technique in an unacceptable way. Thus, apparent clonality is generally related to the specific use or application of the array prepared by the method described herein. Exemplary levels of contamination acceptable in individual amplification sites for a particular application include, but are not limited to, up to 0.1%, 0.5%, 1%, 5%, 10%, or 25% of contaminating amplicons. An array may include one or more amplification sites having these exemplary levels of contaminating amplicons. For example, up to 5%, 10%, 25%, 50%, 75%, or even 100% of the amplification sites in an array may contain contaminated amplicons. It will be understood that in an array or other set of sites, at least 50%, 75%, 80%, 85%, 90%, 95%, or 99% or more of the sites are clonal, or appear to be clonal.

[0150] In some embodiments, binding equilibrium exclusion may occur when a process occurs at a sufficiently fast rate to effectively eliminate the occurrence of another event or process. Take, for example, the preparation of a nucleic acid array in which sites of an array are randomly seeded with indexed fragments from solution, and copies of the indexed fragments are produced in an amplification process to fill each seeded site to capacity. According to the binding equilibrium exclusion method of this disclosure, the seeding and amplification processes can proceed simultaneously under conditions where the amplification rate exceeds the seeding rate. Therefore, the relatively fast rate at which copies are produced at sites seeded with a first target nucleic acid effectively excludes the second nucleic acid from seeding those sites for amplification. The binding equilibrium exclusion amplification method can be implemented as described in detail in the disclosure of U.S. Patent Application Publication No. 2013 / 0338042.

[0151] Binding equilibrium exclusion can take advantage of a relatively slow rate for initiating amplification (e.g., a slow rate for making a first copy of an indexed fragment) versus a relatively fast rate for making subsequent copies of an indexed fragment (or first copies of an indexed fragment). In the example in the previous paragraph, binding equilibrium exclusion occurs due to a relatively slow rate of indexed fragment seeding (e.g., relatively slow diffusion or transport) versus a relatively fast rate at which amplification occurs to fill the site with copies of the indexed fragment species. In another exemplary embodiment, binding equilibrium exclusion may occur due to a delay in the formation of the first copy of the indexed fragment seeded at the site (e.g., delayed or slow activation) versus a relatively fast rate at which subsequent copies are made to fill the site. In this embodiment, several different indexed fragments may be seeded at individual sites (e.g., several indexed fragments may exist at each site before amplification). However, since the formation of the first copy of any given indexed fragment can be activated randomly, the average rate of first copy formation will be relatively slow compared to the rate at which subsequent copies are produced. In this case, individual sites may have several different indexed fragments seeded, but binding equilibrium exclusion allows only one of those indexed fragments to be amplified. More specifically, when the first indexed fragment is activated for amplification, the site is rapidly filled to capacity with its copy, thereby preventing the creation of a copy of the second indexed fragment at the site.

[0152] In one embodiment, the method is carried out simultaneously to (i) transport an indexed fragment to an amplification site at an average transport rate, and (ii) amplify the indexed fragment at the amplification site at an average amplification rate, where the average amplification rate exceeds the average transport rate (U.S. Patent No. 9,169,513). Therefore, in such embodiments, binding equilibrium exclusion can be achieved by using a relatively slow transport rate. For example, since lower concentrations result in slower transport rates, a sufficiently low concentration of the indexed fragment can be selected to achieve the desired average transport rate. Alternatively or additionally, the transport rate can be reduced by using a high-viscosity solution and / or a molecular crowding reagent in the solution. Examples of useful molecular crowding reagents include, but are not limited to, polyethylene glycol (PEG), Ficol, dextran, or polyvinyl alcohol. Exemplary molecular crowding reagents and formulations are described in U.S. Patent No. 7,399,590, which is incorporated herein by reference. Another factor that can be tuned to achieve the desired transport rate is the average size of the target nucleic acid.

[0153] Amplification reagents can contain additional components that promote amplicon formation, and in some cases increase the rate of amplicon formation. One example is a recombinase. Recombinases can promote amplicon formation by enabling repeated infiltration / extension. More specifically, recombinases can promote the infiltration of an indexed fragment by polymerase and the extension of a primer by polymerase, which uses the indexed fragment as a template for amplicon formation. This process can be repeated as a chain reaction in which the amplicon produced from each round of infiltration / extension serves as a template in subsequent rounds. Since no denaturation cycle (e.g., by heating or chemical denaturation) is required, this process can be carried out more rapidly than standard PCR. Therefore, recombinase-enhanced amplification can be carried out isothermally. To promote amplification, it is desirable to include ATP or other nucleotides (or possibly non-hydrolyzable analogs thereof) in the recombinase-enhanced amplification reagent. A mixture of recombinase and single-strand binding (SSB) protein is particularly useful because SSBs can further promote amplification. A representative formulation for recombinase-enhancing amplification is the TwistAmp kit, commercially available from TwistDx (Cambridge, UK). Useful components and reaction conditions for recombinase-enhancing amplification reagents are described in U.S. Patents No. 5,223,414 and No. 7,399,590.

[0154] Another example of a component that can be included in amplification reagents to promote amplicon formation and, in some cases, increase the rate of amplicon formation is helicase. Helicase can promote amplicon formation by enabling a chain reaction of amplicon formation. This process can be carried out more rapidly than standard PCR because no denaturation cycle (e.g., by heating or chemical denaturation) is required. Therefore, helicase-enhanced amplification can be carried out isothermally. A mixture of helicase and single-strand binding (SSB) protein is particularly useful because SSB can further promote amplification. A representative formulation for helicase-enhanced amplification is the IsoAmp kit commercially available from Biohelix (Beverly, Massachusetts). Furthermore, examples of useful formulations containing helicase protein are described in U.S. Patents 7,399,590 and 7,829,284.

[0155] Another example of a component that can be included in amplification reagents to promote amplicon formation and, in some cases, increase the rate of amplicon formation is origin-binding protein.

[0156] Sequencing methods

[0157] Following the attachment of indexed fragments to the surface, the sequences of the immobilized and amplified indexed fragments are determined. Sequencing can be comprehensive sequencing or targeted sequencing. Comprehensive sequencing can be used when the entire sequence of each cell or nucleus present in the library is desired. Examples of applications using comprehensive sequencing include, but are not limited to, whole-genome sequencing, whole-transcriptome sequencing, and ATAC sequencing. Targeted sequencing can be used when information regarding biological characteristics is desired. In one embodiment, targeted sequencing can be used to identify subpopulations of cells or nuclei, or subsets of the genome, transcriptome, proteome, or any combination thereof, as described in detail herein.

[0158] Sequencing can be carried out using any suitable sequencing technique, and methods for determining the sequence of fixed and amplified indexed fragments, such as chain resynthesis, are known in the art and are described, for example, by Bignell et al. (U.S. Patent No. 8,053,192), Gunderson et al. (International Publication No. 2016 / 130704), Shen et al. (U.S. Patent No. 8,895,249), and Pipenburg et al. (U.S. Patent No. 9,309,502).

[0159] The methods described herein can be used in conjunction with various nucleic acid sequencing methods. Particularly applicable techniques involve mounting nucleic acids at fixed positions within an array so that their relative positions do not change, and the array being repeatedly imaged. For example, embodiments in which images are obtained in different color channels corresponding to different labels used to distinguish one nucleotide base type from another are particularly applicable. In some embodiments, the process of determining the nucleotide sequence of an indexed fragment may be an automated process. A preferred embodiment is the synthetic sequencing ("SBS") technique.

[0160] SBS technology generally involves the enzymatic elongation of a nascent nucleic acid chain by the repeated addition of nucleotides to a template chain. In conventional SBS methods, a single nucleotide monomer may be delivered to the target nucleotide in the presence of polymerase at each delivery. However, the method described herein allows for the delivery of multiple types of nucleotide monomers to the target nucleic acid in the presence of polymerase during delivery.

[0161] In one embodiment, the nucleotide monomer comprises locked nucleic acid (LNA) or crosslinked nucleic acid (BNA). The use of LNA or BNA in the nucleotide monomer increases the hybridization intensity between the nucleotide monomer and the sequencing primer sequence present on the immobilized indexed fragment.

[0162] SBS can use nucleotide monomers having a terminator moiety or nucleotide monomers lacking a terminator moiety. Methods using nucleotide monomers without a terminator include, for example, pyrosequencing and sequencing using γ-phosphate-labeled nucleotides, as further described herein. In methods using nucleotide monomers without a terminator, the number of nucleotides added to each cycle is generally variable and depends on the template sequence and the mode of nucleotide delivery. In SBS techniques utilizing nucleotide monomers with a terminator moiety, the terminator may be effectively irreversible under the sequencing conditions used, as in conventional Sanger sequencing using dideoxylinucleotides, or it may be reversible, as in the sequencing method developed by Solexa (now Illumina, Inc.).

[0163] SBS technology can use nucleotide monomers with or without a labeled moiety. Therefore, integration events can be detected based on the characteristics of the label, such as fluorescence; the characteristics of the nucleotide monomer, such as molecular weight or charge; and by-products of nucleotide integration, such as pyrophosphate release. In embodiments where two or more different nucleotides are present in the sequencing reagent, the different nucleotides may be distinguishable from each other, or two or more different labels may be distinguishable under the detection technique used. For example, different nucleotides present in the sequencing reagent may have different labels, and they can be distinguished using a suitable optical system, as exemplified by the sequencing method developed by Solexa (now Illumina).

[0164] A preferred embodiment is pyrosequencing technology. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) when specific nucleotides are incorporated into the nascent DNA chain (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M. and Nyren, P. (1996) "Real-time DNA sequencing using detection of pyrophosphate release." Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001) "Pyrosequencing sheds light on DNA sequencing." Genome Res., 11(1), 3-11; Ronaghi, M., Uhlen, M. and Nyren, P. (1998) "A sequencing method based on real-time pyrophosphate." Science (281(5375), 363, U.S. Patent Nos. 6,210,891, 6,258,568 and 6,274,320). In pyrosequencing, released PPi can be detected by immediate conversion to adenosine triphosphate (ATP) by ATP sulfurase, and the level of the generated ATP is detected via photons produced by luciferase. The nucleic acids to be sequenced can be attached to feature regions in an array, and the array can be imaged to capture the chemiluminescent signal produced by incorporating nucleotides into the feature regions of the array. Images can be obtained after processing the array with a specific nucleotide type (e.g., T, C, or G). The images obtained after the addition of each nucleotide type differ in terms of which feature regions in the array are detected. These differences in the images reflect the different sequence contents of the feature regions on the array. However, the relative position of each feature region remains unchanged in the image. The images can be stored, processed, and analyzed using the methods described herein.For example, images obtained after processing an array with each different nucleotide type can be processed in the same manner as those obtained from different detection channels for a reversible terminator-based sequencing method, as illustrated herein.

[0165] In another exemplary type of SBS, cycle sequencing is achieved by stepwise adding reversible terminator nucleotides containing cleavable or photobleachable dye labels, as described, for example, in International Publication No. 04 / 018497 and U.S. Patent No. 7,057,026. This technique has been commercialized by Solexa (now Illumina) and is also described in International Publication Nos. 91 / 06678 and 07 / 123,744. The availability of fluorescently labeled terminators with reversible ends and cleaved fluorescent labels facilitates efficient cyclic reversible termination (CRT) sequencing. Polymerases can also be co-operated to efficiently incorporate and extend these modified nucleotides.

[0166] In some reversible terminator-based sequencing embodiments, the label does not substantially inhibit extension under SBS reaction conditions. However, the detection label may be removable, for example, by cleavage or decomposition. Images can be captured after the incorporation of the label into the arrayed nucleic acid feature regions. In certain embodiments, each cycle involves the simultaneous delivery of four different nucleotide types to the array, each nucleotide type having a spectrally different label. Four images can then be obtained by using a selective detection channel for each of the four different labels. Alternatively, different nucleotide types can be added sequentially, and an image of the array can be obtained between each addition step. In such embodiments, each image shows a nucleic acid feature region incorporating a particular type of nucleotide. Because the sequence content of each feature region is different, different feature regions may or may not be present in different images. However, the relative positions of the feature regions remain unchanged within the images. Images obtained from such a reversible terminator-SBS method can be stored, processed, and analyzed as described herein. Following the image capture step, the label can be removed, and the reversible terminator portion can be removed for subsequent nucleotide addition and detection cycles. Removing the label after detection in a specific cycle and before subsequent cycles has the advantage of reducing background signal and crosstalk between cycles. Examples of useful labeling and removal methods are described herein.

[0167] In certain embodiments, some or all of the nucleotide monomers may include a reversible terminator. In such embodiments, the reversible terminator / cleavable fluorophore may include a fluorophore attached to the ribose moiety via a 3' ester bond (Metzker, Genome Res. 15:1767-1776 (2005)). Another method involved separating the chemical substance of the terminator from fluorescent labeling (Ruparel et al., Proc Natl Acad Sci USA 102:5932-7 (2005)). Ruparel et al. describe the development of a reversible terminator that uses a small amount of 3' allyl group to block elongation but can be easily unblocked by short-term treatment with a palladium catalyst. The fluorophore was attached to the base via a photocleavable linker that can be readily cleaved by 30 seconds of exposure to long-wavelength UV light. Thus, either disulfide reduction or photocleavage can be used as the cleavable linker. Another technique for reversible termination is the use of a natural termination followed by the placement of a bulky dye on the dNTP. The presence of a charged bulky dye on the dNTP can act as an effective terminator via steric and / or electrostatic hindrance. The presence of one embedded event prevents further binding unless the dye is removed. Cleavage of the dye removes the fluorophore and effectively reverses the termination. Examples of modified nucleotides are also described in U.S. Patents 7,427,673 and 7,057,026.

[0168] Additional exemplary SBS systems and methods that can be used in conjunction with the methods and systems described herein are described in U.S. Patent Publications 2007 / 0166705, 2006 / 0188901, 2006 / 0240439, 2006 / 0281109, 2012 / 0270305, and 2013 / 0260372, U.S. Patent No. 7,057,026, and International Publication No. 05 / 065814, U.S. Patent Publication No. 2005 / 0100900, and International Publication No. 06 / 064199 and 07 / 010,251.

[0169] Some embodiments may utilize the detection of four different nucleotides using fewer than four different labels. For example, SBS can be carried out using the method and system described in U.S. Patent Publication No. 2013 / 0079232, which is incorporated material. As a first example, pairs of nucleotide types can be detected at the same wavelength but may be distinguished based on a difference in intensity for one member of the pair, or based on a change in one member of the pair (e.g., through chemical modification, photochemical modification, or physical modification) that causes a noticeable signal to appear or disappear compared to the signal detected for the other members of the pair. As a second example, three of the four different nucleotide types can be detected under specific conditions, while the fourth nucleotide type has no detectable label under those conditions or is minimally detectable under those conditions (e.g., minimal detection by background fluorescence). Incorporation of the first three nucleotide types into the nucleic acid can be determined based on the presence of their corresponding signals, and incorporation of the fourth nucleotide type into the nucleic acid can be determined based on the absence of any signal or minimal detection. As a third example, one nucleotide type may include a label that is detected by two different channels, while other nucleotide types are detected by one or fewer channels. The three exemplary configurations described above are not considered mutually exclusive and can be used in various combinations.An exemplary embodiment combining all three examples is a fluorescence-based SBS method using a first nucleotide type detected in a first channel (e.g., dATP with a label detected in the first channel when excited by a first excitation wavelength), a second nucleotide type detected in a second channel (e.g., dCTP with a label detected in the second channel when excited by a second excitation wavelength), a third nucleotide type detected in both the first and second channels (e.g., dTTP with at least one label detected in both channels when excited by the first and / or second excitation wavelengths), and an unlabeled fourth nucleotide type (e.g., unlabeled dGTP) that is not detected or minimally detected in any channel.

[0170] Furthermore, as described in the incorporated document, U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single channel. In such a so-called one-dye sequencing method, a first nucleotide type is labeled but the label is removed after the first image is generated, and a second nucleotide type is labeled only after the first image is generated. A third nucleotide type retains its label in both the first and second images, and a fourth nucleotide type remains unlabeled in both images.

[0171] Some embodiments may utilize sequencing by ligation techniques. Such techniques involve incorporating oligonucleotides using DNA ligases and identifying the incorporation of such oligonucleotides. Oligonucleotides typically have different labels that correlate with the identity of specific nucleotides in the sequence into which the oligonucleotide hybridizes. As with other SBS methods, an image can be obtained after processing an array of nucleic acid sequences with a labeled sequencing reagent. Each image shows nucleic acid feature regions incorporating a specific type of label. Because the sequence content of each feature region is different, different images may or may not contain different feature regions, but the relative positions of the feature regions remain constant within the image. Images obtained from ligation-based sequencing methods can be stored, processed, and analyzed as described herein. Exemplary SBS systems and methods that can be used with the methods and systems described herein are described in U.S. Patents 6,969,488, 6,172,218, and 6,306,597.

[0172] In some embodiments, nanopore sequencing can be used (Deamer, DW & Akeson, M. "Nanopores and nucleic acids: prospects for ultrarapid sequencing.", Trends Biotechnol., 18, 147-151 (2000); Deamer, D. and D. Branton, "Characterization of nucleic acids by nanopore analysis," Acc. Chem. Res. 35:817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin, and J. Golovchenko, "DNA molecules and configurations in a solid-state nanopore microscope," Nat. Mater. 2:611-615 (2003)). In such embodiments, indexed fragments pass through nanopores. Nanopores may be synthetic pores such as α-hemolysin or biological membrane proteins. As indexed fragments pass through nanopores, each base pair can be identified by measuring the variation in the electrical conductance of the pores. (U.S. Patent No. 7,001,792; Soni, GV & Meller, "A. Progress toward ultrafast DNA sequencing using solid-state nanopores." Clin. Chem. 53, 1996-2001 (2007); Healy, K., "Nanopore-based single-molecule DNA analysis." Nanomed., 2,459-481 (2007); Cockroft, SL, Chu, J., Amorin, M. & Ghadiri, MR, "A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution." J. Am Chem. Soc. 130, 818-820 (2008).)Data obtained from nanopore sequencing can be stored, processed, and analyzed as described herein. Specifically, the data can be processed as images according to exemplary processing of optical and other images as described herein.

[0173] Some embodiments may utilize methods that include real-time monitoring of DNA polymerase activity. Nucleotide incorporation can be detected, for example, via fluorescence resonance energy transfer (FRET) interactions between fluorophore-containing polymerases and gamma-phosphate-labeled nucleotides, as described in U.S. Patents No. 7,329,492 and 7,211,414; or nucleotide incorporation can be detected using zero-mode waveguides, as described in U.S. Patent No. 7,315,019, and fluorescent nucleotide analogs and manipulated polymerases, as described in U.S. Patent No. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082. Illumination can be restricted to a zeptolite-scale volume around the surface tethering polymerase so that the incorporation of fluorescently labeled nucleotides can be observed with low background (Levene, MJ et al. "Zero-mode waveguides for single-molecule analysis at high concentrations." Science, 299, 682-686 (2003); Lundquist, P et al. "Parallel confocal detection of single molecules in real time." Opt. Lett. 33, 1026-1028 (2008); Levene, MJ et al. "Zero-mode waveguides for single-molecule analysis at high concentrations." Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008)). Images obtained by such methods can be stored, processed, and analyzed as described herein.

[0174] Some SBS embodiments include the detection of protons released during the incorporation of nucleotides into the extension product. For example, sequencing based on the detection of released protons can be performed using electrodetectors and related technologies commercially available from Ion Torrent (Guilford, Connecticut, a subsidiary of Life Technologies), or sequencing methods and systems described in U.S. Patent Applications Publications 2009 / 0026082, 2009 / 0127589, 2010 / 0137143, and 2010 / 0282617. The methods described herein for amplifying target nucleic acids using binding equilibrium exclusion can be readily applied to substrates used for proton detection. More specifically, the methods described herein can be used to produce a clonal population of amplicons used for proton detection.

[0175] The SBS method described above can be advantageously implemented in a multiplex format so that multiple different indexed fragments are manipulated simultaneously. In certain embodiments, different indexed fragments can be processed on a common reaction vessel or on the surface of a specific substrate. This allows for multiplexed delivery of sequencing reagents, removal of unreacted reagents, and detection of embedded events. In embodiments using surface-bound target nucleic acids, the indexed fragments may be in array form. In array form, the indexed fragments can typically be bound to the surface in a spatially distinguishable manner. The indexed fragments can be bound by direct covalent bonding, attachment to beads or other particles, or binding to polymerase or other molecules attached to the surface. The array may contain a single copy of the indexed fragment at each site (also called a feature site), or multiple copies having the same sequence may be present at each site or feature site. Multiple copies can be produced by amplification methods such as bridge amplification or emulsion PCR, which are described in further detail herein.

[0176] The method described in this specification can use, for example, an array having features of any of various densities, including at least about 10 features / cm 2 , 100 features / cm 2 , 500 features / cm 2 , 1,000 features / cm 2 , 5,000 features / cm 2 , 10,000 features / cm 2 , 50,000 features / cm 2 , 100,000 features / cm 2 , 1,000,000 features / cm 2 , 5,000,000 features / cm 2 , or more.

[0177] An advantage of the method described in this specification is that for multiple cm 2The objective is to provide rapid, efficient, and parallel detection. Accordingly, this disclosure provides an integrated system that can prepare and detect nucleic acids using techniques known in the art, such as those exemplified herein. Accordingly, the integrated system of this disclosure may include a fluid component that can deliver amplification reagents and / or sequencing reagents to one or more immobilized indexed fragments, and the system may include components such as pumps, valves, reservoirs, and fluid lines. Flow cells may be configured and / or used in the integrated system for detecting target nucleic acids. Exemplary flow cells are described, for example, in U.S. Patent Application Publication 2010 / 0111768 and U.S. Patent Application 13 / 273,666. As exemplified with respect to flow cells, one or more fluid components of the integrated system may be used in amplification and detection methods. Taking an embodiment of nucleic acid sequencing as an example, one or more fluid components of the integrated system may be used for delivery of sequencing reagents in amplification methods described herein and sequencing methods such as those exemplified above. Alternatively, the integrated system may include separate fluid systems for performing amplification and detection methods. Examples of integrated sequencing systems capable of producing amplified nucleic acids and determining nucleic acid sequences include, but are not limited to, the MiSeq® platform (Illumina, Inc., San Diego, CA) and the apparatus described in U.S. Patent Application No. 13 / 273,666.

[0178] Detection of rare events

[0179] This disclosure also provides methods for identifying and / or characterizing rare events. Currently, methods for characterizing rare events within a population without enrichment are costly and difficult. When enrichment is used, selection is typically based on several biological characteristics of the cell, such as size, morphology, or the presence or absence of identifiable molecules such as proteins or glycans on the cell surface. This limits the types of events that can be identified. The methods presented herein represent a significant advance in the ability to identify and / or characterize the presence or absence of rare events. In general, the present invention provides identification, enrichment, and sequencing-based characterization of subsets of rare single cells present in libraries of millions or billions of cells. Using the identification of rare single cells, a cell database can be created that researchers can use to determine cells available for further analysis.

[0180] Examples of rare events include, but are not limited to, rare cells within a large population of cells. Types of rare cells include, but are not limited to, cell classes, species types, and disease states or risks. Examples of rare cell classes include, but are not limited to, cells from individuals with alterations in the genome, transcriptome, or epigenome. Examples of rare species types include, but are not limited to, prokaryotic cells, eukaryotic cells, or fungal cells. Examples of rare cells associated with disease states or risks include, but are not limited to, cancer cells.

[0181] Rare events are typically identified by biological features (usually the presence or absence of a nucleotide sequence) that correlate with the rare event. In one embodiment, the biological feature is a biomolecule such as a protein, glycan, proteoglycan, or lipid. The biomolecule may be tagged with nucleic acids bound to compounds such as antibodies that specifically bind to the biomolecule. The biological feature may be known in advance (e.g., known before the method is performed and referred to as predetermined) or newly known (e.g., the biological feature is identified after targeted sequencing or comprehensive sequencing as described herein).

[0182] Examples of genome-related biological features include, but are not limited to, modifications in immune cells such as gene rearrangements. Examples of transcriptome-related biological features include the expression of one or more specific genes or RNA molecules, or the expression of specific proteins. Examples of epigenome-related biological features include, but are not limited to, epigenetic patterns such as methylation labels, methylation patterns, and accessible DNA, or the expression of specific proteins that correlate with epigenetic changes. Examples of biological features correlated with rare species types include 16s rRNA or rDNA, 18s rRNA or rDNA, and internal transcriptional spacer (ITS) rRNA / rDNA, or the expression of specific proteins by rare species. Examples of biological features associated with disease status or risk include germline cells or somatic cells having mutant DNA sequences or expression patterns of RNA and / or proteins that correlate with diseases such as cancer.

[0183] This method may include identifying members (individual modified target nucleic acids) of a sequencing library that contain rare events. In one embodiment, this method may include scrutinizing a sequencing library suspected of containing rare events. Scrutinizing a sequencing library typically involves determining (i) biological features that correlate with rare events and (ii) indices present in members of the library for sequences of two types of nucleotide regions present in the library. In one embodiment, sequences of two or more biological features can be determined.

[0184] In one embodiment, the nucleotide sequence of a biological feature is identified by targeted sequencing. Targeted sequencing methods are known in the art and may include the use of primers that hybridize to approach the biological feature in terms of position and orientation, acting as a starting site for sequencing. For example, if the biological feature is the presence or absence of a specific single nucleotide polymorphism (SNP), a primer can be designed to specifically anneal to a nucleotide close to the SNP. In another example, if the biological feature is a protein, a primer can be designed to specifically anneal to a nucleotide of nucleic acid attached to a compound specifically bound to a biomolecule. As a result, those skilled in the art can obtain sequence data that enables the identification of members of a library containing the biological feature of interest. Determining the sequences of indices present in members of a sequencing library is a routine part of single-cell combinatorial indexing.

[0185] Next, sequence data from target sequencing of biological features and sequencing of indices are analyzed using conventional bioinformatics methods to identify these combinations of index sequences present in the same library members as biological features. This correlation between biological features and index sequences identifies subsets of library members, each member including a unique classification of biological features and index sequences, as well as the creation of a cell database. Each unique classification of index sequences, also referred to herein as “marker index sequences,” is similarly present in other members of the library derived from the same cell or nucleus, e.g., the index library under consideration. In one embodiment, the marker index sequences are consecutive indices, i.e., multiple sets of indices present in the library member in rows with 0, 1, 2, 3, 4, or more nucleotides between each index. As described herein, these marker index sequences can be used to focus subsequent sequencing efforts on these members of a library derived from cells or nuclei possessing the biological features, thus reducing costs.

[0186] This method may further include modifying a sequencing library to increase the expression of these members derived from cells or nuclei possessing the biological characteristics. Modification may include enrichment (e.g., positive selection of these rare members in a library containing the desired marker index sequences) or depletion (e.g., negative selection such as the selective removal of rich members in a library that does not contain the desired marker index sequences).

[0187] Enrichment and depletion may involve the use of marker index sequences. Methods for enrichment and depletion are known in the art and include, but are not limited to, marker index sequence-specific amplification (e.g., adapter-immobilized PCR), hybrid capture, and hybridization-based methods such as CRISPR(d)Cas9. Enrichment and depletion methods benefit from the use of nucleotide sequences that specifically hybridize to desired marker index sequences. Thus, enrichment or depletion can be performed on a set of multiple indices present in a library member (see Figure 5B), i.e., sequential indices, i.e., rows having 0, 1, 2, 3, 4 or more nucleotides between each index. Sequential indices that correlate with desired biological features can be reliably selected and retained, thereby enriching the desired library members. Alternatively, sequential indices that do not correlate with desired biological features can be selected and removed, thereby depleting library members that correlate with abundant cells and effectively enriching library members that correlate with the desired biological features. In one embodiment, enrichment may be accompanied by targeted amplification. For example, after constructing a sequencing library, an amplification reaction can be used to specifically amplify library members containing the biological features of interest. In one embodiment, specific amplification can be achieved using a biological feature-specific primer designed to anneal to a nucleotide sequence containing the biological feature, and a second primer to anneal to one side of all members of the library. The biological feature-specific primer may contain one or more indices and / or universal sequences at its 5' end.

[0188] The total length of the sequential index depends on the size of the probe required for specific hybridization between the probe and a member of the library having the desired marker index sequence. In some embodiments, the total length of the sequential index (and therefore the marker index sequence) is at least 40 nucleotides, at least 45 nucleotides, at least 50 nucleotides, or at least 55 nucleotides, and 80 nucleotides or less, 75 nucleotides or less, 70 nucleotides or less, or 65 nucleotides or less. In one embodiment, the total length of the sequential index is 60 nucleotides.

[0189] By using either enrichment or depletion, a sublibrary containing increased expressions of these members of a library derived from cells or nuclei possessing the biological characteristics can be obtained. Comprehensive sequencing of the sublibrary can be performed using conventional methods, such as those described herein. Because the expression is sufficiently increased, comprehensive sequencing requires significantly fewer resources and is therefore cost-effective. By using comprehensive sequencing of the sublibrary, one or more previously unknown biological characteristics can be identified.

[0190] Purpose

[0191] The methods provided herein can be readily incorporated into essentially any application, including the preparation of sequencing libraries such as whole genome, transcriptome, epigenome, accessible (e.g., ATAC), and structural (e.g., HiC). Numerous sequencing library methods are known to those skilled in the art and can be used to construct whole genome or target libraries (see, for example, the “Sequencing Methods Review” available at genomics.umn.edu / downloads / sequencing-methods-review.pdf).

[0192] In these embodiments aimed at detecting rare events, the methods provided by this disclosure can be readily incorporated into essentially any application using single-cell combinatorial indexing (sci) methods, including but not limited to whole-genome (e.g., sci-WGS-seq), epigenomic (e.g., sci-MET-seq), accessible (e.g., sci-ATAC-seq), transcriptome (sci-RNA-seq), and structural (sci-HiC-seq). In some embodiments, the application involves using structural single-cell combinatorial indexing, including proximity ligation using a linked long-read method with crosslinking. In some embodiments, the application is a co-assay, simultaneously evaluating two or more different samples or information from a given sample. Examples of samples include, but are not limited to, DNA, RNA, and proteins (e.g., surface proteins). Examples include assays that analyze the whole genome and transcriptome, or ATAC and transcriptome (Ma et al., 2020, bioRxiv, DOI:doi.org / 10.1016 / j.cell.2020.09.056).

[0193] In some embodiments, the application is metagenomics (the study of genetic material recovered directly from environmental samples). Examples of environments include those found in fields related to agriculture (e.g., soil), biofuels (e.g., microbial communities that convert biomass), biotechnology (e.g., microbial communities that produce bioactive compounds), and gut microbiota (e.g., microbial communities present in the human or animal microbiome). Genetic material may be present in prokaryotic and / or eukaryotic microorganisms (both unicellular and multicellular), such as fungal cells. The methods described herein can be used to identify rare cells, whether or not they can be cultured. Biological features that can be used to identify rare events in metagenomics include, but are not limited to, 16s rRNA or rDNA, 18s rRNA or rDNA, and internal transcription spacer (ITS) rRNA / rDNA, or proteins encoded by microorganisms. After identification, rare cells can be comprehensively sequenced.

[0194] In some embodiments, this application relates to disease status or risk. It can identify rare events such as single nucleotide polymorphisms (SNPs) and / or biomarkers that correlate with disease or the risk of disease, and these cells having SNPs and / or biomarkers are comprehensively sequenced. For example, a fluid biopsy of circulating cells in the bloodstream of a subject, or a tissue biopsy of cells, may be analyzed for rare events related to disease or the risk of disease. Rare events that can be assayed include, but are not limited to, somatic driver mutations that enable the assignment of a particular cancer. A relevant application is to fully characterize and track tumor progression by obtaining samples from a subject over a period of time, selecting cancerous cells or nuclei, and then comprehensively sequencing a subset of tumor cells.

[0195] In some embodiments, this application relates to immune cells. Immune cells undergo rearrangement of specific genes related to the acquired ability of the immune system to identify external molecules. Examples of immune cells that undergo gene rearrangement include, but are not limited to, T cells (e.g., rearrangement of T cell receptors), antigen-presenting cells (e.g., rearrangement of genes encoding major histocompatibility complex proteins), and B cells (e.g., rearrangement of antibodies). Biological features associated with the modification of immune cells may be, but are not limited to, specific rearrangements or proteins derived from specific rearrangements. Immune cells with specific modifications, including, but not limited to, the repertoire characteristics and evolution of T cell receptors, can be fully characterized and tracked. In another embodiment, this application relates to cell differentiation. For example, differentiation events, such as the correlation between accessibility and expression, can be evaluated using expression levels and / or methylation in different regions.

[0196] A non-limiting exemplary embodiment of the present disclosure is shown in Figure 6. In this embodiment, a method for identifying and characterizing the T cell receptor repertoire may include providing a plurality of cells (Figure 6, block 600) and distributing a subset of cells into a plurality of compartments (Figure 6, block 601). The plurality of cells may be, for example, from a blood sample or a lymph node sample. The nucleic acids present in the cells of each compartment are modified by inserting an index (Figure 6, block 602), and then the cells are pooled (Figure 6, block 603). Additional indices are added by a “split and pool” process that repeats distribution (Figure 6, block 601), index addition (Figure 6, block 602), and pooling of subsets (Figure 6, block 603). In one embodiment, each index is added to the same side of the library members to result in a continuous index (see Figure 5B). Optionally, a universal sequence may be added along with one or more indices. After the final index is added, the library of nuclear or intracellular nucleic acids can be pooled (Figure 6, block 603) and further processed to prepare for targeted sequencing of biological features, which enables the identification of T cell receptors containing specific nucleotide sequences, such as those capable of binding to biomolecules of microorganisms or viruses, and the sequencing of indices associated with the biological features of interest (Figure 6, block 604). Sequence analysis (Figure 6, block 605) is used to identify marker index sequences, i.e., intrinsic classifications of index sequences. The identified marker index sequences are those that (i) correlate with biological features and therefore identify members of the library derived from rare cells, or (ii) do not correlate with biological features and therefore identify members of the library derived from abundant cells. A subsequent step in this exemplary embodiment describes the depletion of abundant members of the library, but the method may be modified as described herein to enrich rare library members.Specific oligonucleotide or guide RNA sequences can be designed to hybridize with marker index sequences that correlate with members of a library derived from rich cells (Figure 6, block 606), and the sequencing library of members from rich cells can then be depleted, for example, by using hybridization capture or CRISPR digest (Figure 6, 607). As a result, a modified library is obtained containing an increased expression of these members derived from cells with biological characteristics. Members of the modified sequencing library can be subjected to comprehensive sequencing (Figure 6, block 608). Alternatively, the modified library can be subjected to further enrichment and / or depletion until the expression of the desired members of the library meets the criteria for characterization. For example, members of the modified library can undergo a second sequencing, the marker index is identified, and specific oligonucleotide or guide RNA sequences are designed and used to deplete or enrich the modified library.

[0197] In some embodiments, the application involves the use of a continuous index. A non-limiting exemplary embodiment of an approach to constructing a sequencing library using a continuous index is shown in Figure 7. After the distribution of a subset of cells or nuclei, a first compartment-specific index I1 can be added to DNA molecules 705 present in the cells or nuclei, for example, by tagging (Figure 7, step 701). If the primary source of nucleic acids is RNA, the nucleic acids can be converted to DNA using methods such as cDNA synthesis before tagging. As a result, a library of modified nucleic acids present in the cells or nuclei is obtained, each modified nucleic acid 706 containing a compartment-specific index I1 at each end. The subsets are poolable, and the ends of the resulting modified target nucleic acids can be repaired as needed, for example, by a 3' fill-in. In one embodiment, the 5' end of the modified target nucleic acid may be phosphorylated. In one embodiment, the next step after the second index addition can be facilitated by adding an overhang (e.g., a G, C, or poly-A tail) to the 3' end of the modified target nucleic acid. The pooled cells or nuclei are distributed into a second compartment set, to which a second compartment-specific index I2 may be added, for example, by ligation of an adapter having a appropriately modified 3' end, e.g., a T-tail 3' end (Figure 7, step 702). This yields cells or nuclei containing a library of modified nucleic acids, each modified nucleic acid 707 containing two compartment-specific indices I1 and I2 at each end. The ends of the target nucleic acids can be modified to facilitate the addition of the next index, for example, by phosphorylation of the 5' end and / or modification of the 3' end with a poly-A tail, or by the addition of G or C to the 3' end. If desired, the pooling and addition of other compartment-specific indices can be repeated to add an appropriate number of indices. In one embodiment, when adding the final compartment-specific index I3 to the distributed cells or subset, an adapter having a universal sequence may be included (Figure 7, step 703). For example, a mismatch adapter can be added to each end to obtain modified nucleic acid 708. Examples of universal sequences include those used to immobilize library members into an array (P5 and P7).The mismatch adapter may also contain universal sequences useful for sequencing, or in some embodiments, it may amplify the modified nucleic acid 708 (Figure 7, step 704), and add universal sequences useful for sequencing (i5 and i7) to obtain the modified nucleic acid 709. The modified nucleic acid 709 can be used in target sequencing to identify marker index sequences that correlate with biological features useful for subsequent enrichment and / or deletion.

[0198] Figure 8 shows a non-limiting, exemplary embodiment that combines enrichment with targeted amplification. In this embodiment, a single-cell combinatorial library is prepared (e.g., Figure 3, block 35; Figure 4, block 47; Figure 6, block 605), and the resulting modified nucleic acid (e.g., Figure 7, modified nucleic acid 709) is subjected to an amplification reaction that specifically amplifies library members containing the biological feature of interest. Modified nucleic acid 802 having a continuous index is contacted with a primer 803 that may contain two domains: a 3' domain designed to anneal to a nucleotide sequence containing the biological feature, and one or more universal sequences or their complements, e.g., a 5' domain having i7 and P7. The amplification reaction includes a second primer 804 that anneals to one side of all members of the library. Amplification 801 results in modified nucleic acid 805 having compartment-specific indices I1-3 at one end and a universal sequence added with a two-domain primer targeting the biological feature at the other end. The amplified and modified target nucleic acids can be used in target sequencing and sequencing to identify marker index sequences that correlate with the biological features of interest.

[0199] Kits are also provided herein. In one embodiment, the kit is for preparing a sequencing library. In one embodiment, the kit comprises one transposome complex, including a transposon recognition site so that a universal sequence can be inserted into a target nucleic acid. In another embodiment, the kit comprises two transposome complexes, each complex including a transposon recognition site having a different universal sequence so that the universal sequence can be inserted into a target nucleic acid. In yet another embodiment, the kit comprises components for adding at least one, two, or three indices to a nucleic acid. The kit may also comprise other components useful for constructing a sequencing library. For example, the kit may comprise at least one enzyme that mediates ligation, primer extension, or amplification to process a DNA molecule to include an index. The kit may comprise a nucleic acid having an index sequence.

[0200] The components of the kit are typically contained in suitable packaging material in quantities sufficient for at least one assay or use. Optionally, other components such as buffers and solutions may be included. Typically, instructions for use of the packaged components are also included. As used herein, the term “packaging material” refers to one or more physical structures used to contain the contents of the kit. The packaging material is generally constructed by conventional methods to provide a sterile, contaminant-free environment. The packaging material may have labels indicating that the components can be used to prepare a sequencing library. In addition, the packaging material includes instructions on how to use the materials in the kit. As used herein, the term “package” refers to a container, such as glass, plastic, paper, or foil, that can hold the components of the kit within a certain limit. “Instructions for use” typically include specific descriptions of at least one assay parameter, such as reagent concentration, or the relative amounts of reagent and sample to be mixed, the retention period of the reagent / sample mixture, temperature, and buffering conditions.

[0201] composition

[0202] During or after the preparation of a sequencing library, numerous molecules and compositions may be obtained. For example, the resulting molecules or compositions may contain modified target nucleic acids adjacent to one or both sides by a sequential index. A sequential index may contain 1, 2, 3, 4, 5, 6, or more indices in a row, and each index is separated from other indices by 1, 2, 3, 4, or more nucleotides. In some embodiments, the total length of a sequential index is at least 40 nucleotides, at least 45 nucleotides, at least 50 nucleotides, or at least 55 nucleotides, and 80 nucleotides or less, 75 nucleotides or less, 70 nucleotides or less, or 65 nucleotides or less. Libraries or compositions containing multiple such modified target nucleic acids may be obtained. Pooled libraries and compositions containing pooled libraries of such polynucleotides may be obtained.

[0203] Exemplary Embodiments

[0204] Embodiment 1. A method for identifying a subpopulation of cells containing biological characteristics, (a) To provide a single-cell sequencing library, The sequencing library contains multiple modified target nucleic acids. The modified target nucleic acid contains at least one index sequence, (b) Examining a sequencing library by targeted sequencing to identify index sequences present in modified target nucleic acids that have the same biological characteristics, The index sequences related to biological characteristics are marker index sequences, and (c) Modifying the sequencing library to obtain a sublibrary, The sublibrary contains an enhanced representation of the modified target nucleic acid that includes the marker index sequence, compared to other modified target nucleic acids present in the sequencing library that do not contain the marker index sequence. (d) A method comprising determining the nucleotide sequence of a modified target nucleic acid including a marker index sequence.

[0205] Embodiment 2. The method according to Embodiment 1, wherein the single-cell sequencing library contains nucleic acids from multiple samples.

[0206] Embodiment 3. The method according to any one of Embodiments 1 to 2, wherein the multiple samples include (i) samples of the same tissue obtained from different organisms, (ii) samples of different tissues from one organism, or (iii) samples of different tissues from different organisms.

[0207] Embodiment 4. The method according to any one of Embodiments 1 to 3, wherein in step (b), two or more marker index sequences are identified.

[0208] Embodiment 5. The method according to any one of Embodiments 1 to 4, wherein the single-cell combinatorial sequencing library comprises a target nucleic acid representing the whole genome or a subset of the genome of a cell or nucleus.

[0209] Embodiment 6. The method according to any one of Embodiments 1 to 5, wherein the subset of the genome includes a transcriptome, accessible chromatin, DNA, a three-dimensional structure, or a target nucleic acid representing a cellular or nuclear protein.

[0210] Embodiment 7. A modification of any one of Embodiments 1 to 6, comprising enrichment of a modified target nucleic acid containing a marker index sequence.

[0211] Embodiment 8. The method according to any one of Embodiments 1 to 7, wherein the concentration is a hybridization-based method.

[0212] Embodiment 9. A hybridization-based method is a method according to any one of Embodiments 1 to 8, comprising hybrid capture, amplification, or CRISPR(d)Cas9.

[0213] Embodiment 10. The method according to any one of Embodiments 1 to 9, which is modified to include depletion of a modified target nucleic acid that does not contain a marker index sequence.

[0214] Embodiment 11. A method according to any one of Embodiments 1 to 10, wherein depletion is a hybridization-based method.

[0215] Embodiment 12. A hybridization-based method is a method according to any one of Embodiments 1 to 11, comprising hybrid capture, amplification, or CRISPR(d)Cas9.

[0216] Embodiment 13. The method according to any one of Embodiments 1 to 12, wherein the biological features include a nucleotide sequence indicating the species type.

[0217] Embodiment 14. The method according to any one of Embodiments 1 to 13, wherein the species type includes a cell species.

[0218] Embodiment 15. The method according to any one of Embodiments 1 to 14, wherein the biological features include a nucleotide in the 16s subunit, the 18s subunit, or the ITS non-transcribed region.

[0219] Embodiment 16. The method according to any one of Embodiments 1 to 15, wherein the biological features include a nucleotide sequence indicating a cell class.

[0220] Embodiment 17. The method according to any one of Embodiments 1 to 16, wherein the cell class includes an expression pattern, an epigenetic pattern, immunorecombination, or a combination thereof.

[0221] Embodiment 18. The method according to any one of Embodiments 1 to 17, wherein the epigenetic pattern comprises a methylation label, a methylation pattern, accessible DNA, or a combination thereof.

[0222] Embodiment 19. The method according to any one of Embodiments 1 to 18, wherein the biological features include a nucleotide sequence indicating a disease state or risk.

[0223] Embodiment 20. The method according to any one of Embodiments 1 to 19, wherein the disease state or risk includes a mutated DNA sequence, a mutation expression pattern, or a mutation epigenetic pattern correlated with the disease.

[0224] Embodiment 21. The method according to any one of Embodiments 1 to 20, wherein the mutant DNA sequence includes at least one single nucleotide polymorphism.

[0225] Embodiment 22. The method according to any one of Embodiments 1 to 21, wherein the mutation expression pattern includes the expression of a biomarker.

[0226] Embodiment 23. A mutational epigenetic pattern comprising a methylation label and a methylation pattern, according to any one of Embodiments 1 to 22.

[0227] Embodiment 24. The method according to any one of Embodiments 1 to 23, wherein the modified target nucleic acid comprises consecutive indices of at least two compartment-specific index sequences, and there are no more than seven nucleotides between the two index sequences.

[0228] Embodiment 25. The method according to any one of Embodiments 1 to 24, wherein the continuous index is present at each end of the modified target nucleic acid.

[0229] Embodiment 26. The method according to any one of Embodiments 1 to 25, wherein the length of the continuous index is at least 55 nucleotides.

[0230] Embodiment 27. The method according to any one of Embodiments 1 to 26, wherein one copy of the continuous index is present in the modified target nucleic acid.

[0231] Embodiment 28. The method according to any one of Embodiments 1 to 27, wherein two copies of the continuous index are present in the modified target nucleic acid.

[0232] Embodiment 29. The method according to any one of Embodiments 1 to 28, wherein the multiple modified target nucleic acids in the sequencing library represent at least 100,000 different cells or nuclei.

[0233] Embodiment 30. To provide a single-cell combinatorial sequencing library, The method according to any one of Embodiments 1 to 29, comprising processing a sample to prepare a library, wherein the sample is a metagenomics sample obtained from an organism.

[0234] Embodiment 31. The method according to any one of Embodiments 1 to 30, wherein the organism is a mammal.

[0235] Embodiment 32. The method according to any one of Embodiments 1 to 31, wherein the metagenomics sample includes tissue suspected of containing symbiotic or pathogenic microorganisms.

[0236] Embodiment 33. The method according to any one of Embodiments 1 to 32, wherein the microorganism is a prokaryote or a eukaryote.

[0237] Embodiment 34. The method according to any one of Embodiments 1 to 33, wherein the metagenomics sample includes a microbiome sample.

[0238] Embodiment 35. To provide a single-cell combinatorial sequencing library, A method according to any one of embodiments 1 to 34, comprising processing a sample to prepare a library, wherein the sample is from a living organism.

[0239] Embodiment 36. The method according to any one of Embodiments 1 to 35, wherein the organism is a mammal.

[0240] Embodiment 37. The method according to any one of Embodiments 1 to 36, wherein the primary source of nucleic acids from the sample includes RNA.

[0241] Embodiment 38: The method according to any one of Embodiments 1 to 37, wherein the RNA comprises mRNA.

[0242] Embodiment 39. The method according to any one of Embodiments 1 to 38, wherein the primary source of nucleic acids from the sample includes DNA.

[0243] Embodiment 40. The method according to any one of Embodiments 1 to 39, wherein the DNA comprises whole cell genomic DNA.

[0244] Embodiment 41. The method according to any one of Embodiments 1 to 40, wherein the whole cell genomic DNA comprises nucleosomes.

[0245] Embodiment 42. The method according to any one of Embodiments 1 to 41, wherein the primary source of nucleic acids from the sample includes cell-free DNA.

[0246] Embodiment 43. The method according to any one of Embodiments 1 to 42, wherein the sample contains cancer cells.

[0247] Embodiment 44. The method according to any one of Embodiments 1 to 43, wherein providing a single-cell combinatorial sequencing library is achieved by constructing the library using a single-cell combinatorial indexing method selected from single-nucleus transcriptome sequencing, single-cell transcriptome sequencing, single-cell transcriptome and transposon-accessible chromatin sequencing, single-nucleus whole-genome sequencing, single-nucleus sequencing of transposon-accessible chromatin, single-cell epitope sequencing, sci-HiC, and sci-MET.

[0248] Embodiment 45. The method according to any one of Embodiments 1 to 44, comprising providing two different single-cell combinatorial sequencing libraries from each cell or nucleus.

[0249] Embodiment 46. The method according to any one of Embodiments 1 to 45, wherein two different single-cell combinatorial sequencing libraries are selected from single-nuclear transcriptome sequencing, single-cell transcriptome sequencing, single-cell transcriptome and transposon-accessible chromatin sequencing, single-nuclear whole-genome sequencing, single-nuclear sequencing of transposon-accessible chromatin, sci-HiC, and sci-MET.

[0250] Embodiment 47. The method according to any one of Embodiments 1 to 46, further comprising performing a sequencing procedure to determine the nucleotide sequence of a nucleic acid.

[0251] Embodiment 48. A method for preparing a sequencing library containing nucleic acids from multiple single nuclei or single cells, (a) To provide a plurality of nuclei or cells, wherein the nuclei or cells include nucleosomes, (b) Contacting a plurality of nuclei or cells with a transpososome complex comprising a transposase and a universal sequence, wherein the contact further includes conditions suitable for incorporating the universal sequence into DNA nucleic acid and yielding a double-stranded DNA nucleic acid comprising the universal sequence, (d) Distributing a plurality of nuclei or cells into a first plurality of compartments, Each compartment contains a subset of the nucleus or cells, (e) Processing DNA molecules within each subset of a nucleus or cell in order to generate an indexed nucleus or cell, The process involves adding a first compartment-specific index sequence to the DNA nucleic acids present in each subset of the nucleus or cell, resulting in indexed nucleic acids present in the indexed nucleus or cell. The processing includes ligation, primer extension, hybridization, amplification, or a combination thereof. (g) A method comprising combining indexed nuclei or cells to generate pooled indexed nuclei or cells.

[0252] Embodiment 49. The method according to claim 48, comprising providing a plurality of nuclei or cells in a plurality of compartments, each compartment comprising a subset of nuclei or cells, contacting a transposome complex, and further comprising combining the nuclei or cells after contacting to produce a pooled nucleus or cell.

[0253] Embodiment 50. Provided is the method according to any one of Embodiments 48 to 49, comprising subjecting the nuclei to a chemical treatment in order to produce nucleosome-depleted nuclei while maintaining the integrity of the isolated nuclei.

[0254] Embodiment 51. Distributing a pool of indexed nuclei or cells, including indexed nuclei or cells, into a second set of compartments, Each compartment contains a subset of the nucleus or cells, The process of processing DNA molecules within each subset of a nucleus or cell in order to generate a dual-indexed nucleus or cell, The process involves adding a second compartment-specific index sequence to the DNA nucleic acids present in each subset of the nucleus or cell, resulting in double-indexed nucleic acids present in the indexed nucleus or cell. The processing includes ligation, primer extension, hybridization, amplification, or a combination thereof. The method according to any one of embodiments 48 to 50, further comprising combining double-indexed nuclei or cells to generate pooled double-indexed nuclei or cells.

[0255] Embodiment 52. The method involves distributing pooled nuclei or cells, including dual-indexed nuclei or cells, into a third or more compartments. Each compartment contains a subset of the nucleus or cells, To generate a triple-indexed nucleus or cell, the process involves processing DNA molecules within each subset of the nucleus or cell, The process involves adding a third compartment-specific index sequence to the DNA nucleic acids present in each subset of the nucleus or cell, resulting in triple-indexed nucleic acids present in the indexed nucleus or cell. The processing includes ligation, primer extension, hybridization, amplification, or a combination thereof. The method according to any one of embodiments 48 to 51, further comprising combining triple-indexed nuclei or cells to generate pooled triple-indexed nuclei or cells.

[0256] Embodiment 53. The method according to any one of Embodiments 48 to 52, wherein the dispensing step includes dilution.

[0257] Embodiment 54. The method according to any one of Embodiments 48 to 53, wherein the compartment includes a well, a microfluidic compartment, or a droplet.

[0258] Embodiment 55. The method according to any one of Embodiments 48 to 54, wherein each of the first plurality of compartments contains 50 to 100,000,000 nuclei or cells.

[0259] Embodiment 56. The method according to any one of Embodiments 48 to 55, wherein each of the second plurality of compartments contains 50 to 100,000,000 nuclei or cells.

[0260] Embodiment 57. The compartments of the third plurality of compartments contain 50 to 100,000,000 nuclei or cells, and the method according to any one of Embodiments 48 to 56.

[0261] Embodiment 58. Contacting comprises contacting each subset with two transposome complexes, one transposome complex comprising a first transposase comprising a first universal sequence, and the second transposome complex comprising a second transposase comprising a second universal sequence, and contacting further comprises conditions suitable to incorporate the first universal sequence and the second universal sequence into a DNA nucleic acid to yield a double-stranded DNA nucleic acid comprising the first universal sequence and the second universal sequence, the method according to any one of Embodiments 48 to 57.

[0262] Embodiment 59.Adding a compartment-specific index sequence comprises a two-step process of adding a nucleotide sequence comprising a universal sequence to a nucleic acid and then adding a compartment-specific index sequence to the nucleic acid, the method according to any one of Embodiments 48 to 58.

[0263] Embodiment 60. Further comprising obtaining an indexed nucleic acid from the pooled indexed nuclei or cells, thereby further comprising creating a sequencing library from a plurality of nuclei or cells, the method according to any one of Embodiments 48 to 59.

[0264] Embodiment 61. Further comprising obtaining a dual-indexed nucleic acid from the pooled dual-indexed nuclei or cells, thereby further comprising creating a sequencing library from a plurality of nuclei or cells, the method according to any one of Embodiments 48 to 60.

[0265] Embodiment 62. Further comprising obtaining a triple-indexed nucleic acid from the pooled triple-indexed nuclei or cells, thereby further comprising creating a sequencing library from a plurality of nuclei or cells, the method according to any one of Embodiments 48 to 61.

[0266] Embodiment 63. The process further includes providing a surface containing multiple amplification regions, The amplification site comprises at least two groups of bound single-stranded captured oligonucleotides having a free 3' end. The method according to any one of embodiments 48 to 62, further comprising contacting a surface containing amplification sites with a nucleic acid fragment containing one, two, or three index sequences, under conditions suitable for generating multiple amplification sites, each containing a clonal population of an amplicon, from individual fragments containing multiple indices.

[0267] Embodiment 64. A method for preparing a nucleic acid library, (a) Providing multiple samples, each sample containing multiple cells or nuclei, and each sample containing multiple cells or nuclei residing in one or more separate compartments, (b) Contacting multiple nuclei or cells with a transpososome complex containing a transposase and a universal sequence, provided that the transpososome complex does not contain an index sequence, the contact further includes conditions suitable for incorporating the universal sequence into nucleic acids, (c) Adding a first index sequence to the nucleic acid of each separate compartment, (d) Combining cells or nuclei from separate compartments, (e) Distributing cells or nuclei into multiple compartments, (f) A method comprising adding a second index sequence to the nucleic acids of multiple compartments.

[0268] Embodiment 65. The method according to Embodiment 64, wherein the first index sequence, the second index sequence, or a combination thereof is added by ligation, primer extension, hybridization, amplification, or a combination thereof.

[0269] Embodiment 66. The method according to any one of Embodiments 64 to 65, wherein steps (d) to (e) are repeated to add a third or more index sequences to cells or nuclei in multiple compartments.

[0270] Embodiment 67. The method according to any one of Embodiments 64 to 66, wherein multiple nuclei or cells are fixed.

[0271] Embodiment 68. The method according to any one of Embodiments 64 to 67, further comprising amplification of an indexed nucleic acid after step (c) or step (f).

[0272] Embodiment 69. The method according to any one of Embodiments 64 to 68, further comprising the step (g) of combining nucleic acids from multiple compartments and determining the sequence of the nucleic acids.

[0273] Embodiment 70. The method according to any one of Embodiments 64 to 69, further comprising performing a sequencing procedure to determine the nucleotide sequence of a nucleic acid.

[0274] Embodiment 71. A method for sequencing a single cell or a single nucleus, (a) Uniquely indexing the nucleic acids of each cell or nucleus in the sample, thereby creating an indexed library of each cell or nucleus, (b) Identify one or more indexed libraries of interest from step (a) using biological characteristics, (c) Enriching the target indexed library in step (b) thereby producing an enriched library, (d) A method comprising sequencing the concentrated library from step (c).

[0275] Embodiment 72. The method according to Embodiment 71, wherein the library is derived from cell or nuclear DNA, RNA, or protein.

[0276] Embodiment 73. The method according to any one of Embodiments 64 to 72, wherein the biological feature is DNA, RNA, or protein, or a combination thereof.

[0277] Embodiment 74. The method according to any one of Embodiments 64 to 73, wherein uniquely indexing in step (a) includes associating at least two different indexes with the nucleic acid of the cell or nucleus.

[0278] Embodiment 75. The method according to any one of Embodiments 64 to 74, wherein the at least two different indexes are consecutive indexes.

[0279] Embodiment 76. The method according to any one of Embodiments 64 to 75, wherein the enriched library is produced by positive enrichment.

[0280] Embodiment 77. The method according to any one of Embodiments 64 to at least 76, wherein the positive enrichment includes amplification.

[0281] Embodiment 78. The method according to any one of Embodiments 64 to 77, wherein the positive enrichment includes a capture agent.

[0282] Embodiment 79. The method according to any one of Embodiments 64 to 78, wherein the positive enrichment includes a solid support.

[0283] [[ID=,27]] Embodiment 80. The method according to any one of Embodiments 64 to 79, wherein the enriched library is produced by negative enrichment.

[0284] Embodiment 81. The method according to any one of Embodiments 64 to 80, wherein identifying the target indexed library in step (c) includes sequencing the index.

[0285] Embodiment 82. A method for sequencing a single cell or a single nucleus, comprising: (a) providing a sample, the sample including a plurality of nuclei or cells, (b) Associating a first index with each nucleus or cell in the sample, (c) Dividing the sample into multiple sections, (d) Associating a second index with each nucleus or cell in multiple compartments, (e) Pooling multiple sections, (f) Sequencing the pooled parcels, (g) Identifying combinations of the first and second indices associated with biological characteristics, A method comprising (h) enriching biological features from pooled compartments using identified combinations of a first index and a second index from step (g).

[0286] Embodiment 83. A kit, (a) Multiple transposome complexes, each of which includes a transposase and a transposon sequence, and the transposon sequence is not indexed, (b) A first plurality of index oligonucleotides, wherein the first plurality of index oligonucleotides includes oligonucleotides having at least two different sequences, (c) A kit containing a ligase enzyme for use with an index oligonucleotide.

[0287] Embodiment 84. The kit according to Embodiment 83, further comprising a second plurality of index oligonucleotides, wherein the second plurality of index oligonucleotides comprises oligonucleotides having a different sequence from the first plurality of index oligonucleotides.

[0288] Embodiment 85. The kit according to Embodiment 83 or 84, further comprising a third plurality of index oligonucleotides, wherein the third plurality of index oligonucleotides comprises oligonucleotides having sequences different from those of the first plurality of index oligonucleotides and the second plurality of index oligonucleotides.

[0289] [Examples]

[0290] This disclosure is illustrated by the following examples. It should be understood that specific examples, materials, quantities, and procedures should be interpreted broadly in accordance with the scope and spirit of this disclosure as described herein.

[0291] Example 1

[0292] Human Cell Atlas of Chromatin Accessibility During Development

[0293] summary

[0294] The chromatin landscape of the human genome shapes cell-type specific programs of gene expression. We developed an improved assay for single-cell profiling of chromatin accessibility based on three levels of combinatorial indexing (sci-ATAC-seq3) and applied it to 59 fetal samples representing 15 organs, profiling approximately one million single cells. We annotated this data by leveraging cell types defined by gene expression within the same organ, constructing a catalog of hundreds of thousands of cell-type specific DNA regulatory elements, and investigating the characterization of lineage-specific transcription factors, as well as cell-type specific enrichment of complex traits. Together with the accompanying human cell atlas of gene expression during development, this data constitutes a rich resource for exploring human biology.

[0295] Main text

[0296] In recent years, there has been a surge in single-cell methods, experiments, and atlases. However, the vast majority of these efforts remain focused on single-cell gene expression, reflecting only one aspect of cell biology, developmental biology, and organic biology. Other aspects, such as the chromatin landscape that shapes gene expression programs, are equally important for investigations at single-cell resolution, but they face the challenge of having relatively few scalable methods.

[0297] The single-cell combinatorial indexing ("sci") framework involves splitting cells or nuclei and pooling them into wells, where molecular barcodes are introduced in situ into the target species (e.g., RNA or chromatin) each time. Through sequential, in-situ molecular barcoding, species within the same cell are matched and labeled with a unique combination of barcodes, and sci-assays have been developed to profile chromatin accessibility (sci-ATAC-seq), gene expression (sci-RNA-seq), nuclear structure, genome sequence, methylation, histone labeling, and other phenomena, as well as sci-co-assays to profile chromatin accessibility and gene expression together, for example ("CoBatch," "Split-seq," "Pagaired-seq," and "dscATAC-seq" are also methods that rely on single-cell combinatorial indexing).

[0298] Previously, chromatin accessibility in ~100,000 mammalian cells could be profiled via two levels of sci-ATAC-seq, but the assay has several limitations. For example, it requires custom loading of the Tn5 enzyme with a barcoded adapter, and collisions require 10 per experiment. 4 ~10 5This is limited to individual cells, i.e., cells that accept the same barcode combination. To address these issues, we have developed an improved assay for single-cell profiling of chromatin accessibility based on three levels of combinatorial indexing (sci-ATAC-seq3). In contrast to previous iterations of sci-ATAC-seq, this assay is independent of molecularly barcoded Tn5 complexes (Figure 9; Figure 10). Rather, the first two indexing steps are achieved by ligating either end of a conventional, uniformly packed Tn5 transposase complex (standard "Nextera"), while the final indexing step is still via PCR. Compared to two levels of sci-ATAC-seq, but similar to sci-RNA-seq3, sci-ATAC-seq3 significantly reduces library preparation costs per cell, as well as collision rates. The theoretical collision rates for 2-level indexing (96x384 wells) and 3-level indexing (384x384x384 wells) are 12% and 1.3%, respectively. The observed collision rate for a 3-level "seed mix" experiment using pooled equal numbers of GM12878 and CH12.LX cells is estimated to be 4.0%. 6 This opens the way for cell-scale experiments. This protocol no longer requires cell selection. Furthermore, the inventors optimized the selection of ligases and polymerases, kinase concentrations, and oligo design and concentrations to maximize the number of fragments recovered from each cell. Note that an explicit choice was made to maximize complexity at the expense of specificity of accessible sites while maintaining enrichment within the accessible region. Using Picard, estimated total unique reads ("complexity") were calculated for each cell, and Fraction of Reads in Transcription Start Site ("FRiTSS") was calculated for each cell. Reads within 500 bp of the Gencode TSS were considered to be within the TSS. Specifically, it was found that the sensitivity (i.e., complexity) and specificity (i.e., enrichment at accessible sites) of the assay could be adjusted by adjusting the fixation conditions.

[0299] Towards a human cell atlas of chromatin accessibility, sci-ATAC-seq3 was applied to 59 fetal samples representing 15 organs (adrenal gland, two regions of cerebellum, eye, heart, intestine, kidney, liver, lung, muscle, pancreas, placenta, spleen, stomach, and thymus), and chromatin accessibility was fully profiled in 1.6 million cells (Figure 1D-E). In Example 2, gene expression profiling in 4 to 5 million cells from the same organ is described based on overlapping sample sets. The profiled organs represent a diverse range of systems, with the most noticeable absences being bone marrow, bone, gonads, and skin.

[0300] Rapid and uniform processing of heterogeneous fetal tissue presents a significant challenge. The inventors have developed a novel method for directly extracting nuclei from cryopreserved tissue that functions well across various tissue types and produces homogenates suitable for both sci-ATAC-seq3 and sci-RNA-seq3. In short, rapidly frozen tissue sections are wrapped in aluminum foil and pulverized on dry ice using a cooled hammer. The tissue powder is then divided into aliquots, one for sci-ATAC-seq3 and the other for sci-RNA-seq3.

[0301] For sci-ATAC-seq3, samples were obtained from 23 fetuses with an estimated gestational age ranging from 89 to 125 days. Cells were lysed, nuclei were isolated using the publicly available ATAC-seq cell lysis buffer, and the nuclei were fixed with formaldehyde before rapid freezing for further processing. Approximately 50,000 fixed nuclei from each tissue were deposited across four wells of a 96-well plate and processed for tagging. After tagging, a first index, also identified in the tissue sample, was introduced by ligation into one free end of an asymmetrically inserted transposase complex. After pooling and splitting, a second index was introduced by ligation into the other free end of the transposase complex. Following another round of pooling and splitting, the final index was added by PCR, and the resulting amplicons were pooled for sequencing.

[0302] The sci-ATAC-seq3 library from the third experiment out of five Illumina NovaSeq experiments was sequenced, generating a total of over 50 billion reads. As an initial QC check, the data was examined at the tissue level, i.e., before splitting into single cells. All available single-ended DNase-seq samples from fetal tissue were downloaded from the ENCODE data portal and remapped. Next, accessibility peaks were identified in each "pseudobulk" sample and each ENCODE sample, these were merged, and each sample was scored for accessibility at each peak in the master list. However, while the sci-ATAC-seq3 data were not very concentrated at the peaks (central reads at the peak: 29% for ATAC-SEQ3, 35% for ENCODE DNase-seq), samples from the same tissue correlated similarly for both assays (central spearman correlation: 0.93 for two samples from the same tissue in sci-ATAC-seq3, 0.91 for DNase-seq), and sci-ATAC-seq3 showed higher technical reproducibility (central spearman correlation: 0.95). Furthermore, based on these aggregated profiles, samples were clustered to their respective tissues, either by analyzing sci-ATAC-seq3 samples alone or by analyzing sci-ATAC-seq3 and DNase-seq samples together using pairwise spearman correlation for clustered samples.

[0303] Based on cell barcodes, reads were segmented, and dynamic thresholding was applied as described above to identify 1,568,018 cells. From chicken controls, a collision rate of ~5% was estimated for each of the three experiments. Uniform Manifold Approximation and Projection (UMAP) visualization of cells corresponding to human sentinel tissue did not reveal any obvious experimental batch effects. Three samples were dropped due to nucleosome banding with poor fragment size distribution, and two more samples were dropped because they captured very few cells. In these sci-ATAC-seq3 libraries, it is estimated that the median of 91%–99% of all unique fragments per cell was sequenced for each tissue type.

[0304] After identifying the peak accessibility for each tissue, these were merged to generate a master set of 1.05 million sites. Each cell was scored for the presence or absence of reads at each site, and then low-precision cells were filtered out based on the total number of unique reads (sample-specific minimum in the range of 1,000 to 3,586), the percentage of reads overlapping with the master set of accessible sites (sample-specific minimum in the range of 0.2 to 0.4), the percentage of reads falling near the TSS (+ / -1kb; sample-specific minimum in the range of 0.05 to 0.15), and the doublet score obtained by applying the Scrublet doublet detection algorithm originally developed for scRNA-seq data (excluding ~10% of cells with the highest doublet score).

[0305] Following these procedures, 790,957 single-cell chromatin accessibility profiles remained from 54 fetal samples. The total number of high-precision cells per tissue ranged from 2,421 (spleen) to 211,450 (liver). The median number of unique fragments per cell in this set was 6,042, the median overlap with the master set of accessible sites was 0.49, and the percentage around TSS (+ / -1kb) was 0.19.

[0306] The inventors subjected high-precision cells to latent semantic indexing (LSI) for each tissue using logarithmically transformed term frequency components. Although no clear evidence of batch effects for different samples corresponding to the same tissue was observed, the Harmony algorithm was applied to conserve the selection of samples in the PCA space for each tissue. Using the tissue-aligned PCA space, Louvain clustering was then applied to initially obtain 172 clusters across all tissues. The dimensionality of each tissue dataset was further reduced using UMAP.

[0307] Cell type annotation

[0308] As demonstrated by the inventors and others, cell type annotation in scATAC-seq datasets can be significantly simplified by leveraging scRNA-seq datasets. To partially automate cell type annotation for scATAC-seq data, we first annotated cell types in scRNA-seq data for the same tissue, as described in the manual. Secondly, we calculated gene-level accessibility scores for the scATAC-seq data and aggregated the number of transposition events that fit into the gene body extended by 2kb upstream of their TSSs. Thirdly, we used a gene-cell matrix for each data type as input to an approach to find possible correspondences between scRNA-seq clusters and scATAC-seq clusters based on non-negative least squares (NNLS) regression, thereby obtaining an initial "liftover" set for automated annotation of scATAC-seq clusters. Finally, we manually reviewed all automated annotations by examining the pile-up around marker genes for each cell type within each tissue and corrected assigned labels where deemed necessary. First, cell types were annotated using sci-RNA-seq data collected from matching tissues based on marker gene expression. Louvain clusters were identified using ATAC data for each tissue. Next, gene-level accessibility scores were calculated for each of these clusters and matched to RNA clusters based on non-negative least squares (NNLS) regression, resulting in some merging of Louvain clusters. These initial automated annotations were further refined by manually reviewing the cluster-specific accessibility landscape around marker genes. The annotated cell types showed specific accessibility around the TSS of known marker genes. For each cell type or unannotated cluster, the accessibility around the TSS of known marker genes was summed and the scale was normalized to account for the difference in total reads per cell, as well as the total number of cells across the cell type.The data suggested that some unannotated clusters may not represent novel cell types, but rather technical artifacts (e.g., doublets). While the inventors noted that other approaches have shown great promise for multimodal incorporation of single-cell data, they found that the cluster-versus-cluster NNLS method is sufficient and far less computationally intensive for the purposes of this specification.

[0309] In total, we were able to annotate 150 (87%) of the 172 clusters, or 163 (95%) of the 172 clusters if unreliable labels were included. Some clusters received identical annotations within the same tissue and were therefore merged, resulting in 124 annotations across all tissues. Of these, some annotations were present across multiple tissues (e.g., erythroblasts in 4 tissues). By rejecting across tissues, we obtained 54 (or 59 if unreliable labels and 1:2 mappings) unique cell type annotations that mapped 1:1 to the annotations made on the scRNA-seq dataset. Many of the ScRNA-seq cell types that were not found in the chromatin-accessible data at this level of resolution are small clusters that may not have been sufficiently sampled to be detectable due to the small number of cells profiled in this study (~4M (RNA) vs. ~800K (ATAC) high-precision cells). On the other hand, the majority of the nine scATAC-seq clusters that remained completely unannotated are thought to be due to unfiltered doublets. This is because UMAP representation is characterized by accessibility to marker genes of multiple adjacent cell types.

[0310] Identification of lineage-specific TFs

[0311] Next, we attempted to integrate and compare chromatin accessibility across cell types across all 15 organs. To mitigate the effects of net differences in cell count per organ and / or cell type, 800 cells per cell type were randomly sampled for each organ (or all cells were taken if fewer than 800 cells of a given cell type were shown in a given organ), and UMAP visualization was performed. For reassurance, cell types were shown not in batches or individually, but clustered together across multiple organs, such as stromal cells (9 organs), endothelial cells (13 organs), lymphoid cells (7 organs), and bone marrow cells (10 organs). Cell types related to development and function, such as diverse hematopoietic cells, secretory cells, PNS neurons, and CNS neurons, were also co-localized.

[0312] A key question in developmental biology is which transcription factors (TFs) are involved in producing this diverse range of cell types from an immutable genome. Next, we sought to leverage the breadth of this human cell atlas of chromatin accessibility to systematically evaluate differentially accessible TF motifs and thus identify key regulators of cell fate in the context of in vivo human development.

[0313] As a first approach, we used a linear regression model to identify the TF motifs found in the accessible sites of each cell that best explain the relationships between cell types. We first treated each tissue independently and identified the most highly enriched motifs / TFs from the JASPAR database in each of the 124 annotated cell type clusters to reveal both known and potentially novel regulators. For example, in the placenta, the SPI1 / PU1 motif (an established regulator of myeloid growth) was highly enriched in the myeloid cell peak, the TWIST-1 motif (required for stromal progenitor cell formation) was enriched in the stromal cell peak, and the FOS::JUN motif was associated with chromatin accessibility in the extravillous trophoblast (a cell type in which the corresponding AP1 complex is described as specifically active).

[0314] Interestingly, unannotated clusters within the placenta were highly enriched for the GATA1::TAL1 motif (an established regulator of erythropoiesis). These cells clustered with erythroblasts from other tissues within the global UMAP, and upon further examination, major erythrocyte marker genes showed specific promoter accessibility. This cluster was not annotated in the NNLS-inducible workflow. This is because erythroblast clusters were not detected in the placenta in scRNA-seq studies, presumably because the placenta is one of the few tissues that have ATACs more than RNA cells. Therefore, motif enrichment can aid in cell type annotation when the major regulators of a cell type are known.

[0315] The inventors repeated this analysis for 54 major cell types observed across all tissues, i.e., after discarding cell types appearing in multiple tissues. As expected, the top motifs remained consistent with tissue-specific analyses and literature, e.g., SPI1 / PU1 in bone marrow cells, CRX in retinal pigment and photoreceptor cells, MEF2B(31) in cardiomyocytes and skeletal muscle cells, and SRF in endocardial and smooth muscle cells. Most motifs were enriched in only one or two cell types, while neuronal TF motifs such as OLIG2, NEUROG1, and POU4F1 were enriched in multiple neuronal cell types. Another notable exception was HNF1B, conventionally associated with kidney and pancreatic development, whose motif was enriched in 13 cell types across a range of specific epithelial and secretory cells.

[0316] POU2F1 is an example of a TF that has not been previously associated with a specific developmental branch, but rather is an exception within the POU family, suggesting that it is widely expressed and does not control any particular orbital. In contrast, we have found that its motif is enriched in several neuronal cell types, at least during human fetal development. Further supporting this, POU2F1 is specifically expressed in those same cell types.

[0317] Extending these observations, we then utilized a companion scRNA-seq atlas to more generally examine whether TFs are differentially expressed in patterns consistent with the differential accessibility of their motifs. For example, looking at all cell types annotated in the same tissue in both datasets, the expression of the bone marrow precursor factor SPI1 / PU1 strongly correlates with the enrichment of its motif at accessible sites. Interestingly, this analysis also revealed many TFs that had a negative correlation between expression and motif enrichment. Further examination revealed that these TFs tended to be repressors. For example, GFI1B is described as acting as a crucial repressor in erythroblast and megakaryocyte development by supplementing histone deacetylases upon motif binding, inducing chromatin closure, for example, at the fetal hemoglobin locus. Consistent with this, we observed that its expression negatively correlated with the enrichment of its motif at accessible sites.

[0318] The inventors have found that when TFs are classified as "activators" or "inhibitors" based on GO terms, TF expression and motif accessibility tend to correlate positively with annotated activators and negatively with annotated inhibitors, and that the correlation between motif enrichment and expression can be used to predict the mode of action of unclassified TFs. Exceptions can be largely explained by the absence or competition of GO terms, but a literature search will place them into the category predicted by the correlation values. Therefore, this type of analysis can provide a systematic approach to classifying TFs as activators or inhibitors. For example, NFATc3 is generally described as an activator, but our analysis shows that it exhibits an inhibitory mode of action, particularly in the development of T cells that are highly expressed but whose motifs are depleted at accessible sites. Such an inhibitory mode of action of NFATc3 has been suggested in previous literature. Apart from general classification, insights can also be gained into the context of cell types in which TFs may variably act as activators or inhibitors. For example, it has been shown that TFs such as FOXO3 act as activators in their unmodified state, but act as repressors when phosphorylated. This may explain the more ambiguous relationship between expression and accessibility.

[0319] The approach described above allows for the systematic association of known TFs with potentially novel roles, has the advantage of not relying on the pre-selection of differentially accessible sites for each cell type, and has the further advantage of being able to associate TF expression with the accessibility of its corresponding motif. However, it is limited in that it relies on a database of known TF motifs. As a different approach, specificity scores were also calculated for each accessible site, and 2,000 of the most specific peaks were selected for each cell type and compared with CpG-matched background genome sequences to search for newly enriched motifs within this set. In general, the top novel motifs for individual cell types coincide with the top known motifs identified by linear regression. Interestingly, some cell types that did not have a strong match for known motifs (e.g., endothelial cells, stromal cells, Schwann cells) were still strongly associated with novel motifs. Such results, particularly for endothelial cells, are described further below.

[0320] Cross-sectional analysis of blood cells and endothelial cells

[0321] The nature of this dataset creates an opportunity to investigate organ-specific differences in chromatin accessibility within a wide range of cell types, such as hematopoietic cells and endothelial cells. In the first pass of hematopoietic cell type annotation, we were able to distinguish between myeloid cells, lymphocytes, erythroblasts, megakaryocytes, and hematopoietic stem cells. By extracting these hematopoietic lineages from all organs and reclustering them, it became possible to additionally identify macrophages, B cells, NK / ILC3 cells, T cells, and dendritic cells, again employing an RNA-assisted annotation approach (notably, analyzing similar cell types from multiple tissues requires an additional doublet washing step; see "Methods"). Macrophages can be further classified into groups associated with tissue origin, as previously observed, as well as phagocytic macrophages. This latter group was identified primarily in the spleen, followed by the liver and adrenal glands. Of particular interest among the hematopoietic lineages are erythroblasts, which are due to the spatiotemporal dynamics of erythrocyte formation during fetal development. The inventors first detected this lineage in the liver, adrenal gland, heart, and placenta, and further identified erythroblasts in the superficially profiled spleen (initially annotated only with megakaryocytes and bone marrow cells) through cross-sectional tissue analysis. The proportion of erythroblasts within the hematopoietic lineage in the tissues was highest in the liver (consistent with this organ being the main site of erythrocyte production at this developmental stage), followed by the spleen and adrenal gland, mimicking the trends observed in RNA data. The unexpected result of the adrenal gland being a potential site of fetal hematopoiesis will be further discussed in Example 2.

[0322] Further investigation of erythroblasts revealed that proximal regions of both adult β-globulin and fetal γ-globulin genes are accessible at this developmental stage, while the promoter of the embryonic ε-globulin gene is inaccessible. The erythroblast cluster can be further subdivided into five major Louvain clusters with differential chromatin accessibility, including separate erythroblast progenitor cell clusters. Accessible sites within the erythroblast progenitor cell cluster, as well as adjacent early erythroblast clusters (erythroblast_3), are enriched for GATA1:TAL1 and other GATA motifs. By comparing the expression levels of various GATA factors in erythroblast progenitor cells, GATA1 / 2 can be designated as a likely TF involved in this motif enrichment. Other erythroblast clusters corresponding to later stages of erythrogenesis show motif enrichment of NFE2 / NFE2L2 (erythroblast_1) and KLF factors (erythroblast_2 / 4), and it is noteworthy that enrichment of GATA motif accessibility is conspicuously absent. Recent published scRNA-seq studies on the mouse hematopoietic system reported that GATA2 is induced early in erythropoiesis, and although GATA2 expression subsequently decreases, GATA1 remains stably expressed. In contrast, studies of a selected bulk human in vitro culture population revealed a decrease in GATA1 expression from progenitor cells to differentiated erythroblasts (consistent with observations in human fetal tissue), as well as an increase in KLF1 and NFE-2 levels in late-stage erythroblasts. These results further suggest the possibility of epigenetically distinct subpopulations of differentiated erythroblasts whose accessibility landscape is shaped by non-GATA factors such as KLF1 or NFE-2. For example, the upstream distal regulatory element of GYPA, used as an erythrocyte entry receptor by the malaria parasite, is most accessible in erythroblast_1 and contains a motif similar to the NFE-2 motif.

[0323] Another intriguing cross-tissue system is the vascular endothelium. Interestingly, no TFs are described as being exclusively expressed in vascular endothelial cells, suggesting that the endothelium-specific transcriptome is combinatorially regulated by several TFs that have overlapping expression in the endothelium. Consistent with this, analysis of JASPAR motifs did not reveal any strong enrichment in endothelial cells. On the other hand, the discovery of novel motifs at 2,000 of the most endothelium-specific peaks revealed strong enrichment across background genomic sequences of motifs similar to ERG and SOX15. These motifs tended not to be given much weight in our linear modeling approach because they are not limited to endothelial cells (ERG motifs are more enriched in megakaryocytes, and SOX15 is enriched in several cell types), and the expression of these TFs is not limited to these cell types. Therefore, ERG, while already described as a major regulator of endothelial function, also promotes culture transition to megakaryocytes.

[0324] Endothelial cells are present in all organs and are necessary for both structural and highly differentiated functions, such as gas exchange in the lungs or fluid filtration in the kidneys. In this study, endothelial cells were detected from 13 of 15 organs (with the exception of the cerebellum and eye, which were profiled more superficially). When these cells were extracted across organs and reclustered, they separated significantly according to tissue origin, in contrast to the erythroid lineage, despite a rigorous repeated filtration process to remove any residual contaminant doublets. This allowed us to observe a tissue-specific program of gene expression, as described in Example 2. Indeed, the closest accessibility peaks to these differentially expressed genes had higher specificity scores in the tissues matched in the ATAC data. Furthermore, endothelial cells from almost all organs showed enrichment of specific TF motifs. Notably, many of the enriched TF motifs were differentially expressed in the tissues matched in the RNA data.

[0325] Overall, these findings demonstrate that the general program of chromatin accessibility and gene expression in endothelial cells, a widely distributed cell type that must satisfy both general and organ-specific functions, is mediated by a combination of structural TFs such as ERG and SOX15, as well as tissue-specific TFs that facilitate further specialization. These analyses also highlight the merits of combining novel motif enrichment at specific peaks with a linear modeling approach across tissues, naming key regulators underlying the chromatin accessibility landscape of individual cell types.

[0326] Another interesting example involves the PAEP_MECOM-positive cell type in the placenta, identified in both the scRNA-seq atlas and the sc-ATAC-seq atlas. The regulatory region of this lineage is strongly enriched for the HNF1B motif, a factor traditionally associated with kidney and pancreatic development. For example, HNF1B is expressed very specifically in the PAEP_MECOM cell lineage within the placenta. The nature of ATAC-seq data, which captures some genomic reads across chromosomes even in inaccessible regions, allows for sex differentiation of cells based on Y chromosome or autosomal reads on the X chromosome. Interestingly, we found that PAEP_MECOM and IGFBP1_DKK-positive placental cell types, and to a lesser extent placental bone marrow cells, have a significantly lower Y chromosome read ratio in male fetuses. Consistent with what is known about PAEP (Glycodel) and IGFBP1, these cell types may correspond to maternal endometrial epithelium and stromal cells, respectively.

[0327] CICERO

[0328] As a resource for further research, the inventors generated Cicero co-accessibility scores and Cicero gene activity scores for each tissue in the dataset. The Cicero co-accessibility scores can be used to predict cis-regulatory interactions between accessible elements. The inventors combined elements paired by positive co-accessibility scores to create a database of predicted cis-regulatory interactions. This database contains 80 million unique co-accessible pairs, including 4.5 million (6%) promoter-end pairs, 76 million (94%) end-end pairs, and 128,000 (0.2%) promoter-promoter pairs. The inventors found an average of 33 million co-accessible pairs per tissue. 38% of pairs were specific to a single tissue only, and only 0.007% of pairs were detected in all 16 tissues. Pairs detected in more tissues were more likely to be promoter-end and promoter-promoter pairs. The generated accessibility score and gene activity score can be downloaded from the inventors' website.

[0329] Notably, compared to a control set of 2,040 cells (120 cells randomly selected from each of 17 samples; see supplementary materials), 89% of the 436,206 sites initially identified were significantly differentially accessible (DA), with a false discovery rate (FDR) of 1% in at least one of these 85 cell clusters. To identify DA sites whose accessibility was restricted to specific clusters, metrics used to quantify gene expression specificity in scRNA-seq studies were adapted to chromatin accessibility and calculated for all 436,206 sites across all 85 clusters. 39% of the accessible sites (167,981 / 436,206) were classified as cluster-limited (i.e., increased accessibility in a limited number of clusters), and 55% of these (92,334 / 167,981) were limited to a single cluster.

[0330] Suggestions for cell types in common human traits and diseases

[0331] The majority of the heritability of common human traits and diseases, as measured by genome-wide association studies (GWAS), is fragmented into terminal regulatory elements that are often cell type-specific. Consequently, much of the research is spent cross-referencing GWAS signals with bulk DNase hypersensitivity data (and other epigenetic features) with the aim of systematically linking specific diseases to specific tissue dysfunctions. However, the elucidatement of such studies is significantly limited by cell type heterogeneity. Given the degree of preservation of chromatin accessibility between mice and humans, we considered whether the data could be used to further understand the cell type-specific effects of various genes underlying complex human traits, regardless of interspecies differences. Therefore, despite the fact that our data were generated in mouse tissues, we sought to apply state-of-the-art methods to detect cell type-specific enrichment of human heritability.

[0332] To do this, we used split linkage disequilibrium (LD) score regression (LDSC) to quantify the enrichment of human trait heritability within DA peaks for each of the 85 clusters. After translating human SNPs to orthologous coordinates of the mouse genome, we calculated the enrichment of heritability for 32 phenotypes across the DA peaks obtained for each of the 85 clusters. Of the 85 cell types, 55 had enrichment for at least one phenotype, and of the 32 phenotypes, 28 were enriched for at least one cell type. As a major trend, we observed strong enrichment of heritability for autoimmune diseases such as lupus, celiac disease, and Crohn's disease within the clusters corresponding to leukocytes, while enrichment occurred in neuronal cell types for neurological traits such as bipolar disorder, educational attainment, and schizophrenia. Notably, most of these enrichments were not prominent in the peaks from bulk tissue, demonstrating the values ​​of cell types defined by single-cell chromatin accessibility data. Many enrichments were as expected. For example, the highest concentrations of heritability for low-density lipoprotein (LDL) cholesterol, high-density lipoprotein (HDL) cholesterol, and triglycerides are found in hepatocytes, but interestingly, LDL cholesterol was also significant in renal epithelium of the loop of Henle. Similarly, the highest concentration of heritability for immunoglobulin A (IgA) deficiency is found within T cell clusters. These signals can also lead to a further understanding of the importance of cell subtypes. As an example of this trend, the concentration of heritability for bipolar disorder has been observed in multiple neuronal clusters, but the highest concentration is associated with excitable neurons. In contrast, heritability for Alzheimer's disease is not concentrated in any class of neurons. Instead, its highest concentration is found in microglia clusters.

[0333] To extend our analysis to a larger set of traits, we downloaded GWAS summary statistics for 2,419 traits from over 300,000 individuals from the UK Biobank (nealelab.github.io / UKBB_ldsc / ). Focusing on 405 traits with an effective sample size ≥ 5,000 and estimated heritability ≥ 0.01, we observed a significant enrichment of heritability in 273 traits with at least one cell type, while 74 out of 85 cell types showed enriched heritability for at least one trait. For autoimmune and neurological traits, the same large trends described above are observed here, but the far greater number of traits measured by the UK Biobank reveal further trends. For example, numerous measurements of body size and composition (e.g., body mass index) are also associated with cell types in the brain (Figure 18B). In addition, specific subsets of T cells (12.1, 12.2) are more strongly associated with asthma and allergic rhinitis than other cell types, such as other T cell clusters. At a more granular level, heart attacks are associated with endothelial cells from the liver (25.3), but not with endothelial cells from other endothelial clusters. On the other hand, gout is associated with renal proximal tubular cells. The framework demonstrated herein can be readily applied to single-cell chromatin accessibility data collected from any human or mouse tissue and any heritable trait.

[0334] One result of the new design is compatibility with both 2-level ("2lv2," i.e., "2-level version 2 protocol") and 3-level ("3lv2") configurations, providing greater flexibility in test designs (Figure 9).

[0335] Finally, various conditions for fixing cells or nuclei with formaldehyde were tested to enable long-term stable storage. The inventors found that the choice of buffer used for fixation, as well as the isolation of nuclei before or after fixation, presents a choice between complexity and specificity. In the current study, the inventors select a fixation protocol that increases complexity / sensitivity at the expense of specificity, which can be determined by the end user of the protocol.

[0336] Materials and methods

[0337] cell culture

[0338] Gm12878 cells were cultured and maintained in RPMI 1640 medium (Thermo Fisher Scientific catalog number 11875-093) containing 15% FBS (Thermo Fisher catalog number SH30071.03) and 1% Pen-strep (Thermo Fisher catalog number 15140122). These were counted and divided three times per week at a density of 300,000 cells / mL. The CH12-LX mouse cell line was provided by Michael Snyder lab (Stanford). The cells were cultured in RPMI 1640 medium containing 10% FBS, 1% Pen-strep (penicillin and streptomycin) and 1 × 10^5 M B-ME. These were counted and maintained at a density of 1 × 10^5 cells / mL, and divided three times per week to maintain cell concentration. Both cell lines were incubated at 5% CO2 and 37°C.

[0339] Nuclear isolation and fixation from cell lines

[0340] For cell suspensions, obtain approximately 10 to 100 million cells and pelletize them by rotating at 500xg at room temperature for 5 minutes. Aspirate the supernatant and resuspend the pellet in 1 mL of Omni-ATAC lysis buffer (10 mM NaCl, 3 mM MgCl2, 10 mM Tris-HCl pH 7.4, 0.1% NP40, 0.1% Tween20, and 0.01% digitonin) and incubate on ice for 3 minutes. Add 0.1% Tween20 to 5 mL of 10 mM NaCl, 3 mM MgCl2, and 10 mM Tris-HCl pH 7.4 and pelletize at 500xg at 4°C for 5 minutes. Aspirate the supernatant and resuspend the nuclei in 5 mL of 1X DPBS (Thermo Fisher catalog number 14190144). To crosslink the nuclei, 140 μL of 37% formaldehyde was added in one step to methanol (VWR catalog no. MK501602) to a final concentration of 1%. The fixed mixture was incubated at room temperature for 10 minutes, inverted every 1-2 minutes. To quench the crosslinking reaction, 250 μL of 2.5 M glycine was added and incubated at room temperature for 5 minutes, then incubated on ice for 15 minutes to completely stop the crosslinking. 20 μL of the quenched crosslinked mixture was placed in 20 μL of trypan blue for counting. The crosslinked nuclei were rotated at 500xg at 4°C for 5 minutes, and the supernatant was aspirated. The immobilized nuclei were resuspended in an appropriate amount of freeze buffer (pH 8.0, 50 mM Tris, 25% glycerol, 5 mM Mg(OAc)2, 0.1 mM EDTA, 5 mM DTT (Sigma-Aldrich catalog no. 646563-10X0.5 mL), and a 1× protease inhibitor cocktail (Sigma-Aldrich catalog no. P8340) to obtain 2 million nuclei per 1 mL aliquot, rapidly frozen in liquid nitrogen, and stored at -80°C.

[0341] Organizational procurement and storage

[0342] The target tissue was isolated, washed with 1X HBSS (containing Ca. and Mg.), and then absorbed and dried on a semi-moistened gauze. The dried tissue was placed on a sturdy foil or in a cryotube and rapidly frozen using liquid nitrogen. The frozen tissue was stored at -80°C.

[0343] Nuclear isolation and fixation of frozen fetal tissue

[0344] On the day of grinding, place a cloth towel between the dry ice and the metal to pre-cool the pre-labeled tubes and hammer on the dry ice. Prepare a "packing" using 18-inch x 18-inch sturdy foil, fold it in half twice to make a rectangle, and then fold it two more times to make a square. Place the frozen tissue inside the foil "packing," and then place the tissue in the foil packing inside a pre-cooled 4mm plastic bag to prevent the tissue from falling onto the dry ice if the foil bursts. Cool this tissue packet between two sheets of dry ice. Manually grind the tissue inside the packet using a pre-cooled hammer. After 3-5 impacts, stop the grinding motion and take a break to prevent the sample from overheating. Cool the hammer as needed and repeat the grinding until the tissue is uniform. Divide the ground tissue equally into pre-labeled, pre-cooled 1.5 mL LoBind and nuclease-free snap-cap 1.5 mL tubes (Eppendorf catalog number 022431021). Aliquots of the powdered material can be stored at -80°C until further processing is needed.

[0345] On the day of nucleus isolation, add the lysis buffer directly to the tube, or place the frozen aliquots in a 60 mm dish containing cell lysis buffer and further subdivide using a blade. Unless the aliquots thaw at some point during storage, the powdered tissue aliquots should be easily withdrawn from the storage tube without sample loss. We estimate approximately 20,000 cells per 1 mg of initial tissue weight, and performance may vary from tissue to tissue. Resuspend the pulverized tissue in 1 mL of Omni lysis (RSB + 0.1% Tween + 0.1% NP-40 and 0.01% digitonin), then transfer to a 15 mL Falcon tube. Incubate the nuclei on ice for 3 minutes, then add 5 mL of RSB + 0.1% Tween 20. Centrifuge the nuclei at 500 × g at 4°C for 5 minutes. Aspirate the supernatant and resuspend in 5 mL of 1X DPBS. Remove tissue clumps by passing nuclei in 1X DPBS through a 100 micron cell strainer (VWR catalog number 10199-658). Crosslink the nuclei in a fume hood by adding 140 μL of 37% formaldehyde to methanol in one step to a final concentration of 1%, and quickly mixing by inverting the tube several times. Incubate at room temperature for exactly 10 minutes, gently inverting the tube every 1-2 minutes. Quench the crosslinking reaction by adding 250 μL of 2.5 M glycine (freshly prepared and filter-sterilized), and mix well by inverting the tube several times. Incubate at room temperature for 5 minutes, then incubate on ice for 15 minutes to completely stop crosslinking. Count the nuclei using a hemocytometer to confirm the final amount of freezing buffer to add. The goal is to freeze ~1-2 million nuclei / tube. The cross-linked nuclei are centrifuged at 500xg at 4°C for 5 minutes, the supernatant is aspirated, and the pellet is resuspended in 1-10 mL of freeze buffer supplemented with 1x protease inhibitor and 5 mM DTT. The nuclei are rapidly frozen in liquid nitrogen and stored at -80°C.

[0346] Processing of SCI ATAC-Seq3 samples (library construction and QC)

[0347] Remove the frozen fixed nuclei from -80°C and place them on a bed of dry ice. Thaw the nuclei in a 37°C water bath until thawed (~30 seconds~1 minute), then transfer them to a 15 mL Falcon tube. Pellet the nuclei at 500xg at 4°C for 5 minutes. Aspirate the supernatant without disturbing the pellet, resuspend the pellet in 200 μL of Omni lysis buffer, and then incubate on ice for 3 minutes. Wash the lysis buffer with 1 mL of ATAC-RSB containing 0.1% Tween20, and mix by gently inverting the tube three times. Take 20 μL of nuclei and 20 μL of trypan blue and count the nuclei. Keep the nuclei on ice as much as possible while counting. In a 3-level indexing experiment at 384^3d, the nuclear input number is 4.8 million nuclei per well per tissue @ 50,000 nuclei, or the sample diffused over 96 reactions. The nuclei were pelleted and resuspended in a pre-prepared tagging reaction master mix (Nextera TD buffer, 1X DPBS, 0.1% digitonin, 0.1% Tween 20, and water). Using a wide-mouth tip (Rainin Instrument Co catalog number 30389249), 47.5 μL of nuclei in the tagging mix were divided equally across a LoBind 96-well plate (Eppendorf catalog number 30129512). 2.5 μL of Nextera v2 enzyme (Illumina Inc catalog number FC-121-1031) was added per well, the plate was sealed with adhesive tape, and rotated at 500xg for 30 seconds. The plate was incubated at 55°C for 30 minutes to tag the DNA. The tagging reaction was stopped by adding 50 μL of stop reaction mixture (40 mM EDTA with 1 mM spermidine), and then incubated at 37°C for 15 minutes. Using a wide-mouth tip, tagged nuclei were pooled, pelletized at 500xg at 4°C for 5 minutes, and then washed with ATAC-RSB containing 0.1% Tween 20. The nuclei were pelleted at 500xg at 4°C for 5 minutes, the supernatant was aspirated, and the nuclei were resuspended in 384 μL of ATAC-RSB containing 0.1% Tween 20.A PNK reaction master mix was prepared using 1X PNK buffer (NEB catalog number M0201L), 1 mM rATP (NEB catalog number P0756S), water, and T4 polynucleotide kinase (NEB catalog number M0201L), and added to the nuclei. 5 μL of the PNK reaction mix was divided equally into four LoBind 96-well plates, sealed with adhesive tape, and rotated at 500xg, 4°C for 5 minutes. The PNK reaction was incubated at 37°C for 30 minutes. 13.8 μL of ligation master mix was prepared using 1X T7 ligase buffer (NEB, catalog number M0318L), 9 μM N5 sprint (IDT), water, and 2.5 μL of T7 DNA ligase enzyme (NEB catalog number M0318L) is added directly to the PNK reaction. Using a multi-channel, i.e., 96-head dispenser (Liquidator, catalog number 17010335), 1.2 μL of 50 μM N5 oligo (IDT) is added to each well across four 96-well plates. The plates are sealed with adhesive tape and rotated at 500 x g for 30 seconds, then incubated at 25°C for 1 hour. After the initial ligation, 20 μL of 40 mM EDTA containing 1 mM spermidine is added to stop the ligation reaction, and incubated at 37°C for 15 minutes. Using a wide-mouth tip, each well is pooled into a trough and transferred to a 50 mL Falcon tube. The nuclei are pelleted at 500 x g at 4°C for 5 minutes, the supernatant is aspirated, and 0.1% tween is used. Resuspend the nuclei in 1 mL of ATAC-RSB containing 20 to wash away all residual ligation reaction mix. Pellet the nuclei at 500xg at 4°C for 5 minutes, and aspirate the supernatant without disturbing the pellet. Prepare the N7 ligation master mix (1X T7 ligase buffer, 9 μM N7_sprint (IDT), water, and T7 DNA ligase) and resuspend the nuclei in the ligation master mix. Transfer the nuclei suspended in the master mix to a trough, and using a wide-mouth tip, divide 18.8 μL of the ligation master mix equally into four 96-well LoBind plates, then add 1.2 μL of 50 μM N7_oligo (IDT) to each well across the four 96-well plates.Seal the plate with adhesive tape, rotate at 500xg for 30 seconds, then incubate at 25°C for 1 hour, then add 20 μL of 40 mM EDTA and ImM spermidine to stop ligation, and incubate at 37°C for 15 minutes. Pool the wells into a trough using a wide-mouth tip, then transfer to a 50 mL Falcon tube. Pellet the nuclei at 500xg at 4°C for 5 minutes, aspirate the supernatant, and resuspend the nuclei in 2 mL of Qiagen EB buffer (Qiagen catalog no. 19086). Obtain 20 μL of resuspended nuclei and 20 μL of trypan blue and count the nuclei. Dilute the nuclei to 100-300 nuclei / μL and divide 10 μL / well equally into four 96-well LoBind plates. To reverse crosslink the nuclei, a reverse crosslinking master mix was prepared using EB buffer, proteinase k (Qiagen, catalog no. 19133), and 1% SDS (1 μL / 0.5 μL / 0.5 μL / well), and 2 μL was added to the nuclei in each well. The plates were sealed with adhesive tape, rotated at 500xg for 30 seconds, and incubated at 65°C for 16 hours. Test PCR amplification was performed, and the reaction with SYBR green was monitored in several wells of the plate to determine the optimal number of cycles. Based on the test PCR results, the remainder of the reverse crosslinked plates was amplified with 7.5 μL of NPM, 0.5 μL of BSA (NEB, catalog no. B9000S), 1.25 μL of indexed P5_10 μM (IDT), 1.25 μL of indexed P7_10 μM (IDT), and water per well. Depending on the tissue and nucleus recovery after two ligations, 11–13 cycles were typical for the inventors. The cycle conditions were 3 minutes at 72°C, 30 seconds at 98°C, or 10 seconds at 98°C, 30 seconds at 63°C, and 1 minute at 72°C for 11–13 cycles, followed by a hold at 10°C. The amplification products from the 96-well plates were pooled in a trough, purified using Zymo Clean&Concentrate-5 (Zymo Research catalog number D4014) according to the manufacturer's specifications, and divided into four columns. Each column was eluted in 25 μL of EB buffer and then combined into a single tube.100 μL of AMPure beads (Agencourt, catalog number A63882) were added to the purified PCR product to further remove all residual primer dimers, following the manufacturer's purification process. The final library was eluted from the beads in 25 μL of Qiagen EB buffer. The final library was quantified using a D5000 ScreenTape (Agilent catalog numbers 5067-5588 ScreenTape, 5067-5589 reagent) and an Agilent 4200 Tapestation System to measure the nM concentration of fragments that cluster the wells during sequencing, establishing a 200–1000 base pair window. A 2 nM pool was prepared from equimolar pooling and sequenced at a loading concentration of 1.8 pM using a NextSeq high-power 150-cycle kit (Illumina catalog number 20024904) with a custom recipe and primers.

[0348] Data processing for method development

[0349] Data processing for chicken experiments conducted to develop sci-ATAC-seq3 was performed as described above. Briefly, BCL files were converted to fastq files using bcl2fastq v2.16 (Illumina). Each read was associated with a cell barcode consisting of four components, with a row address appended for tagging and PCR at the P5 end of the molecule and a column address appended for tagging and PCR at the P7 end of the molecule. To correct errors in these barcodes, the inventors divided them into these four components and corrected them to the nearest barcode within an edit distance of 2, provided that the correction was unique at the required edit distance. If none of the four barcodes could be corrected to a known barcode, the corresponding read pair was dropped. The reads were then trimmed with Trimmomatic using the option "ILLUMINACLIP:{adapters_path}:2:30:10:1:true TRAILING:3 SLIDINGWINDOW:4:10 MINLEN:20". Next, the modified reads were mapped to the hybrid human / mouse (hg19 / mm9) gene using bowtie2 with the option "-X 2000-3 1". Subsequently, reads that did not map to a suitable pair on the genome with at least 10 accuracy were filtered out using samtools with the option "-f3-F12-q10", and only reads that mapped to autosomes or sex chromosomes were retained for downstream analysis. Read deduplication was performed per cell barcode using a custom script. Note that, unlike the tissue pipeline (discussed below), read pairs are not maintained redundantly.

[0350] Data processing for tissue samples

[0351] The method for processing sequencing data from tissue samples faithfully follows the methods used and has many optimizations to scale to larger datasets, but for convenience, it is described herein. BCL files were converted to fastq files using bcl2fastq v2.20 (Illumina). Reads with modified barcodes included in the read name were written to separate R1 / R2 files for each sample in our dataset. All mismatch mappings to a known set of barcodes were pre-calculated (feasible because the barcode lengths are short and relatively few in number), and the modification script was executed using pypy (an alternative to the cpython interpreter which is extremely fast for this particular task), and this calculation was parallelized across different lanes of the sequencing run. This resulted in an overall improvement in runtime that significantly outperformed previous methods.

[0352] Next, using the option "ILLUMINACLIP:{adapters_path}TRAILING:3 SLIDINGWINDOW:4:10 MINLEN:20", Trimmomatic was used to adjust the low-precision base / adapter sequences from the 3' end. Then, using bowtie2 with the option "-X 2000 3 1", the adjusted reads were mapped to the hg19 reference genome. Read pairs that did not uniquely map to an autosome or sex chromosome with at least 10 mapping precision were then filtered out using Samtools--samtools view-L{whitelist of chromosomes}-f3-F12-q10-bS. The resulting BAM files were sorted, the aligned reads for each sample were merged using sambabamba, and the resulting BAM files were indexed. This process was parallelized across samples / lanes as much as possible, but runtime would be improved by increasing the number of threads per process by providing trimmomatic / bowtie2 / sambabamba.

[0353] Next, we identified intracellular PCR duplication by identifying a unique set of fragment endpoints within each cell. In our previous research, the resulting duplicate BAM files did not always maintain correct read names between read pairs written to the duplicate BAM files (representative R1 and R2 reads were independently and randomly selected for each unique fragment), which caused compatibility issues with some tools such as SnapATAC (github.com / r3fang / SnapATAC). We corrected this problem and also wrote files that strictly mirror 1) a BED file of cell-specific fragment endpoints and 2) the fragments.tsv.gz file provided by 10x Genomics for scATAC solutions.

[0354] Within each sample, the BED file of cell-specific fragment endpoints was used to recall peaks for each sample using MACS2--macs2 callpeak-t{bed}-f BED-g hs--nomodel--shift-100--extsize 200--keep-dup all--call-summits-n{sample_name}-o{output_dir}. The resulting {outdir} / {sample_name}_peaks.narrowPeak file was sorted and output as a BED file. Peak recalls from all samples included in the downstream analysis (additionally excluding our own standards) were merged using bedtool to form a master set of peaks. As previously described, the use of BED files for peak recall in this specification is intentional and should be noted as it does not take into account the behavior of macs2 in relation to BAM input. When a BAM file is used as input, MACS2 either discards one of the read pairs that uses R1 / R2 independently (effectively downsampling the input data), or, if the BAM file is explicitly specified to be an end pair, uses the entire insert during coverage calculation (the inventors prefer to calculate coverage only for endpoints, not along the entire insert). By using a BED file, coverage can be calculated using all the data and only a window around the molecular endpoints.

[0355] Furthermore, for each sample, a sparse matrix was created counting 1) reads that fall into the peak master set and 2) reads that fall into the gene body extended by 2kb upstream and the 5kb genomic window. In addition, the total number of reads for each cell from the annotated TSS (+ / -1kb around each TSS), the ENCODE blacklist region, and the peak set merged for QC purposes were listed.

[0356] Furthermore, peaks were constructed using a motif matrix, employing the method used in the 10x Genomics scATAC pipeline (see support.10xgenomics.com / single-cell-atac / software / pipelines / latest / algorithms / overview). Briefly, the 10x method allows for the independent discovery of motif occurrences within each bin by calculating the GC% distribution of peaks and bin peaks into equiquartile ranges of GC content. The MOODS package is used to identify background nucleotide compositions matched to each GC bin for motif occurrence and GC bias for motifs in the JASPAR motif database at a p-value threshold of 1E-7. These hits are used to construct motifs by a peak matrix, which can be used to calculate the motif matrix by cell number in downstream analyses. This matrix is ​​binarized so that only one instance of the motif can be counted per peak.

[0357] Cellular barcodes were isolated from background barcode distributions using a modified version of the method used in the 10x genomics scATAC pipeline (see link above). Briefly, this involves fitting a mixture of two negative binary (noise vs. signal) components. Instead of the method used by 10x, k-means was applied to a log-scaled total fragment count distribution to establish an initial threshold between these two distributions, obtaining the maximum cluster with a lower mean total count as the initial threshold. This initial threshold was used to determine the starting parameters of the two distributions using maximum likelihood estimates and further refined by an expectation maximization approach. As described in 10x, this fit can be improved by applying a left shift to the count distribution. Unlike the 10x method, this shift was determined by trying several shifts from 2 to 12 to obtain a mixed distribution model with the best fit. Finally, in contrast to the 10x approach, this method is applied to the distribution of total fragment counts rather than the distribution of counts within the called peaks. The final thresholds selected are the smallest number that yields odds ratios of 20 or greater (beneficial to the signal), and remove at least 0.5% of the signal distribution as estimated from the CDF of the signal distribution (we found that this second criterion hinders fitting with thresholds that otherwise appear overly ambiguous).

[0358] Cellular-level QC, dimensionality reduction, and clustering

[0359] As described above, the total number of unique reads that fall around the TSS (+ / 1kb) in the peak and ENCODE blacklist region was tabled for each cell. Using these totals, for each sample, a sample-specific cutoff for the percentage of unique reads in the peak and the percentage of unique reads falling into the TSS, as well as a global cutoff of 0.5% of unique reads obtained from the ENCODE blacklist region, was selected by visual inspection of these distributions. For a small number of samples with significantly lower autothresholds than other samples in the dataset, a global threshold of 1000 unique reads per cell (or 500 unique fragments per cell) was applied to increase the autothreshold for the corresponding samples. Nucleosome banding scores developed previously were examined, but these scores were not used in QC because no clear distribution of outliers was observed, as previously observed for mouse testes. Before downstream processes, peaks that overlapped with the ENCODE blacklist region or corresponded to sex chromosomes were removed (the latter to avoid introducing potential batch effects between samples of different sexes). Additionally, peaks exceeding two standard deviations from the mean of the log-scale count per peak distribution were excluded, thereby removing peaks with very low counts within the analyzed tissues.

[0360] All downstream processes performed one tissue at a time by pooling cells passed through the entire sample of a given tissue.

[0361] After filtering, a modified version of the Scrublet algorithm was used to remove the cells most likely to be doublets. Briefly, doublets are simulated as the sum of randomly selected cells from the dataset using peaks based on the cell matrix. Next, an LSI is performed using the original cell matrix and the simulated doublets as described below. Note that in this step, similar to how Scrublet applies a multiplier from the original dataset for scRNA-seq data, inverse document frequency (IDF) terms obtained from the original dataset are used instead of simulated doublets. The nearest neighbors of each cell are found in the resulting 50-dimensional space, and the proportion of pseudo-doublets in the neighborhood is calculated as the doublet score. The top 10% of cells in each sample with the highest doublet scores are excluded.

[0362] Regarding dimensionality reduction, we initially found that performing latent semantic indexing (LSI; in other words, latent unanalyzed, i.e., LSA), as previously described, did not work well with the data collected in this study. We suspected this might be due to sparseness and investigated several alternative methods, including CisTopic and SnapATAC. Each of these methods initially appeared to work better than LSI. Initially, the reason for this situation remained unclear, even considering the fundamental similarities of these methods and the nature of the data. We discovered that a simple logarithmic scaling of term frequency terms in LSI, which had not been done before, yielded performance very similar to the other tools tested. This may be due to the exponential distribution of total counts per cell and the strong outlier influence on the PCA process in LSI without logarithmic scaling. This is detailed at andrewjohnhill.com / blog / 2019 / 05 / 06 / dimensionality-reduction-for-scatac-data / . It should be noted that the observed difference with and without the use of logarithmic scaling is particularly large in sparse datasets where the range of total counts per cell is large. Also note that other groups have favorably compared LSI to all other existing methods for reducing the dimensionality of scATAC, confirming our independent findings. Furthermore, we chose to use peaks, as we primarily did in previous studies, because we observed very similar performance when using either genome peaks or a 5kb window.

[0363] In summary, at some point, LSI was performed on one tissue at a time with a binarized window using the cell matrix from all traversing cells of each tissue. First, all parts of individual cells were weighted logarithmically (total number of accessible peaks within the cell) (logarithmically scaled "term frequency"). These weights were then multiplied logarithmically (1 + inverse frequency of each part of all cells), i.e., "inverse document frequency". Next, singular value decomposition was used on the TF-IDF matrix to generate a lower dimensional representation (PCA) of the data, retaining only the 2nd to 50th dimensions (as the 1st dimension tends to correlate highly with read depth). Then, L2 normalization was performed on the PCA matrix to further account for differences in the number of unique fragments per cell. This L2-normalized PCA matrix was used for all downstream processes.

[0364] Although we did not observe significant evidence of batch effects between samples, we applied the Harmonary batch correction algorithm to the PCA space to compensate for batch effects between different samples. Harmony was chosen primarily because it can be easily scaled to large datasets and allows the use of existing PCA coordinates.

[0365] This corrected L2-normalized PCA space was used as input to Louvain clustering and UMAP, as performed in Seurat V3.

[0366] Specificity score

[0367] Before calculating the specificity score, all peaks overlapping with the ENCODE blacklist region were filtered out. The specificity score was calculated for each site / cell type pair as described above.

[0368] Concentration of motifs

[0369] Before calculating motif enrichment, all peaks overlapping with the ENCODE blacklist region were filtered out. First, the motif × cell matrix is ​​obtained by multiplying the peak × motif matrix by the corresponding peak × cell matrix (summed across all cells in the subset of the target data, as described above). Note that the dataset is downsampled to include a maximum of 800 cells per annotation (e.g., cell type) to reduce computational cost and mitigate the over-appearance of very large numbers of cell types during enrichment calculations in downstream processes. Next, for each annotation, negative binomial regression is performed using the speedglm package to predict the total motif count, using two input variables: the annotation indicator column as the principal variable and the logarithm per cell (total number of non-zero entries in the input peak matrix) as the covariate. The coefficients and intercepts of the annotation indicator column are used to estimate the multiplier change in the motif count of the target annotation relative to cells from all other annotations, i.e., exp(intercept+annotation_efficient) / exp(intercept). This test is performed on all motifs in all groups, and then the p-values ​​are corrected using the Benjamini-Hochberg procedure.

[0370] Example 2

[0371] Human Cell Atlas of Gene Expression in Development

[0372] summary

[0373] The emergence and differentiation of cell types during human development are fundamentally fascinating. A single-cell profiling assay for gene expression based on three levels of combinatorial indexing (sci-RNA-seq3) was applied to 121 fetal tissues representing 15 organs, profiling transcription in a total of 4-5 million single cells. From these data, cell types are identified and annotated in relation to marker genes, expression, and regulatory modules. Initial analysis of these data focuses on cell types spanning multiple organ systems, e.g., epithelial cells, endothelial cells, and blood cells. Interesting observations include organ-specific endothelial specialization, potentially novel sites in fetal erythrocytes, and potentially novel cell types. Combined with an accompanying human cell atlas of chromatin accessibility during development, these data represent a rich resource for exploring human biology.

[0374] Main text

[0375] For several reasons, we embarked on generating a human cell atlas that encompasses both gene expression and chromatin accessibility using tissues obtained during development. First, genetic disorders, largely involving developmental components, account for a highly disproportionate proportion of childhood morbidity and mortality. These include thousands of Mendelian disorders, in which both genetic and non-genetic factors contribute significantly, as well as more common diseases (e.g., congenital heart failure, other birth defects, neurodevelopmental disorders, etc.). A reference cell atlas generated from tissue development can serve as a foundation for systematic efforts to understand the specific molecular and cellular events that contribute to each of these childhood diseases.

[0376] Secondly, developing tissues offer a far better opportunity than adult tissues to study the in vivo emergence and differentiation of human cell types. Compared to embryonic and fetal tissues, adult tissues are occupied by differentiated cells and do not simply represent many cellular states. Due to the better resolution of in vivo developmental trajectories, single-cell atlases generated from developing tissues can widely disseminate a fundamental understanding of in vivo human biology, as well as our fundamental understanding of cell reprogramming and cell therapy.

[0377] Thirdly, while pioneering cell atlases have already been reported for many adult human organs, the independent nature of these studies makes it difficult to investigate the differences between cell types appearing in different tissues, such as epithelial cells, endothelial cells, and blood cells. Specifically, comparisons based on existing data are difficult due to differences in sample processing and technological platforms among the groups generating organ-specific cell atlases.

[0378] Towards a human cell atlas of gene expression, a recently developed assay for single-cell RNA-seq, based on three levels of combinatorial indexing (sci RNA-seq3), was applied to 121 fetal tissues representing 15 organs, profiling gene expression in a total of 5 million cells (Figure 11). Example 1 describes profiling of chromatin accessibility in 1.6 million cells from the same organ, based on overlapping sample sets. The profiled organs represent a diverse range of systems, with the most noticeable absences being bone marrow, bone, gonads, and skin.

[0379] Samples were obtained from 28 fetuses with an estimated gestational age ranging from 72 to 129 days. Briefly, these were rapidly frozen, pulverized, and the resulting powder was divided for different assays. For sci-RNA-seq3, nuclei were extracted directly from the cold, fused powder and then fixed with paraformaldehyde. In the kidney and digestive organs, which are rich in RNases and proteases, cells fixed with paraformaldehyde rather than nuclei were used to increase cell and mRNA recovery. For each experiment, nuclei or cells from a given tissue were deposited in different wells, thereby identifying the source as the first index of the sci-RNA-seq3 protocol. As a batch control for experiments with nuclei, a mixture of human HEK293T and mouse NIH / 3T3 nuclei, or nuclei from common "sentinel" tissues (also used in sci-ATAC-seq3 experiments) were placed in one or more wells. As a batch control for cell experiments, cells derived from typical pancreatic tissue (with the nuclei also profiled) were placed in one or more wells.

[0380] Sci-RNA-seq3 libraries were sequenced from seven experiments across seven Illumina NovaSeq runs, generating a total of 68.6 billion reads. The data were processed as described above, and 4,979,593 single-cell gene expression profiles (UMI > 250) were recovered. Single-cell transcriptomes from human-mouse control wells were overwhelmingly species-coherent (~5% collision). Uniform Manifold Approximation and Projection (UMAP) of nuclei or cells from sentinel tissue showed that cell type differences outweighed batch effects between any two experiments. Integrated nuclear and cell analyses corresponding to common pancreatic tissue using Seurat also yielded highly overlapping distributions.

[0381] The inventors profiled a median of 72,241 cells or nuclei per organ (maximum 2,005,512 (cerebrum), minimum 12,611 (thymus)). Compared to other large single-cell RNA-seq atlases, a comparable number of UMIs (median 863 UMIs and 525 genes) were recovered per cell or nucleus, despite relatively shallow sequencing (~14,000 raw reads per cell). As expected, the nucleus showed a higher rate of UMIS mapping to introns than cells (56% for nuclei, 45% for cells, p<2.2e-16, double-sided Wilcoxon rank-sum test). Unless otherwise specified, "cell" is used to refer to both cells and nuclei.

[0382] Tissues were readily identified as originating from either male (n=14) or female (n=14) based on the expression of sex-specific genes. Each of the 15 organs was represented by multiple samples (median 8), including at least two of each sex and a range of gestational age. UMAP visualization of the “pseudobulk” transcriptome of each tissue, clustered by organ rather than individually or experimentally. Approximately half of the expressed protein-coding transcripts were differentially expressed across this set of pseudobulk transcriptomes (11,766 out of 20,033, FDR 5%).

[0383] Scrublet was applied to detect 6.4% of putative doublet cells, corresponding to a 12.6% doublet estimate that included both intra-cluster and inter-cluster doublets. A previously developed strategy for a 2 million mouse organogenic cell atlas (MOCA) was then applied to remove low-precision cells, tablet-enriched clusters, and spike-in HEK293T and NIH / 3T3 cells. All analyses described below are based on 4,062,980 human single-cell gene expression profiles derived from 112 fetal tissues remaining after this filtering process.

[0384] Identification of 77 major cell types

[0385] After filtering for low-precision cells and doublet-enriched clusters, 4 million single-cell gene expression profiles were subjected to UMAP visualization and organ-based Louvain clustering using Monocle 3. Overall, 172 cell types were initially identified and annotated based on cell type-specific markers from the literature. Rejecting annotations common to the tissue reduced the number to 77 major cell types, of which 54 were observed in a single organ (e.g., Purkinje neurons in the cerebellum) and 23 were observed in multiple organs (e.g., vascular endothelial cells in each organ). These 77 major cell types contained a median of 4,829 cells, ranging from 1,258,818 cells (endoexcitatory neurons in the cerebrum) to just 68 cells (SLC26A4_PAEP-positive cells in the adrenal gland). Each major cell type contributed to multiple individuals (median 9). The inventors recovered nearly all major cell types identified by previous atlas-making efforts targeting the same organ, despite differences in species, developmental stage, and technology. A median of 12 major cell types was identified for each organ, ranging from 5 (thymus) to 16 (eye, heart, and stomach). No correlation was observed between the number of cells profiled and the number of cell types identified (ρ=-0.10, p=0.74).

[0386] On average, we identified 11 marker genes per major cell type (minimum 0, maximum 294; differentially expressed genes were defined as those with at least a 5-fold difference in expression between the first and second-ranked cell types; FDR 5%). Several cell types lacked marker genes at this threshold due to similar cell types in other organs (e.g., ENS glial and Schwann cells). Therefore, we also reported sets of “tissue-derived marker genes” determined for each organ using the same procedure (an average of 147 markers per cell type; minimum 12, maximum 778).

[0387] While canonical markers are commonly observed and were indeed important in this annotation process, to our knowledge, the majority of the markers observed are novel. For example, OLR1, SIGLEC10, and the non-coding RNA RP11-480C22.1, along with more established microglial markers such as CLEC7A, TLR7, and CCL3, are among the strongest microglial markers. As a prediction based on the assumption that these tissues are actively growing, many of the 77 major cell types include states that progress from precursor to one or more terminally differentiated cell types. For example, brain excitatory neurons show a continuous trajectory from PAX6+ neuronal precursors to NEUROD6+ differentiated neurons, and further to SLC17A7+ mature neurons. In the liver, hepatic precursors (DLK1+, KRT8+, KRT18+) show a continuous trajectory to functional hepatoblasts (SLC22A25+, ASS2+, ASS1+). In contrast to organogenesis in mice, where the maturation of the transcriptional program is tightly linked to developmental time, cellular state orbitals consistently correlated with estimated gestational age in these human data. The simplest explanation is that gene expression is significantly more dynamic during the early stages of development (i.e., organogenesis vs. fetal development). However, heterogeneous representation and inaccuracies in estimated gestational age may also complicate our understanding.

[0388] In addition to manual annotation of these cell types, we used Garnett to generate semi-automatic classifiers for each organ, as well as a global classifier. The Garnett classifiers were generated independently of clustering using marker genes compiled individually from the literature. The Garnett classifications were remarkably consistent with manual classifications; for example, 88% of cells in the pancreas matched (cluster expansion; 5% mismatch; 7% unclassified). Using Garnett models trained on this human cell atlas, it was also possible to accurately classify cell types from other single-cell datasets, including data from different methods and from adult organs. For example, we applied the Garnett classifier for pancreas to inDrop single-cell RNA-seq data and found that this model accurately annotated 82% of the cells (cluster expansion; 11% inaccurate; 8% unclassified). These Garnett models are posted on our website and can be widely used for automated classification of single-cell data from diverse organs.

[0389] Integration across tissues and investigation of unexpected cell types

[0390] Next, we attempted to integrate the data across all 15 organs and compare cell types. To mitigate the effects of net differences in the number of cells sampled per organ and / or cell type, we randomly sampled 5,000 cells per cell type for each organ (or, if fewer than 5,000 cells of a given cell type were shown in a given organ, all cells were taken), and performed UMAP visualization based on the most differentially expressed genes across cell types within each organ. As expected, cell types shown in multiple organs, such as stromal cells, lymphoid endothelial cells, and mesodermal cells, were generally clustered together. Developmental cell types, such as diverse blood cells, PNS neurons, and mesenchymal cells, were also generally co-localized.

[0391] Using this global UMAP, we identified cell types in organs that were not initially observed and were not clearly annotated or were not expected. In many cases, co-localization with cell types annotated by global UMAP revealed their identity. For example, observing cells in the lungs and adrenal glands that are highly correlated with trophoblast megacells from the placenta (e.g., expressing high levels of placental lactogen, chorionic gonadotropin, and aromatase) suggests that these are trophoblast cells (CSH1_CSH2 positive cells) that have entered the fetal circulation. More surprisingly, we observed placental and spleen cells (AFP_ALB_ positive cells) that are highly correlated with hepatoblasts (e.g., expressing high levels of serum albumin, α-fetoprotein, and apolipoprotein).

[0392] In the heart, we observed three cell types that were not anticipated based on previous atlas-making efforts. The first of these (SATB2_LRRC7-positive neurons) strongly correlates with CNS excitatory neurons and expresses markers including SATB2, PTPRD, and DAB1. To our knowledge, this is an unexpected observation. While we cannot completely rule out contamination from other tissues, we observed these cells at a consistent rate (range) in each sampled heart (n=9), and furthermore, we did not observe any other CNS-like cell types within the heart. The other two are highly correlated with cardiomyocytes but express distinct programs that may reflect specific roles. Specifically, ELF3_AGBL2-positive cardiomyocyte-like cells specifically express many genes associated with alveolar surfactant-secreting cells, such as pulmonary secretory protein 1 (SCGB3A2), pulmonary surfactant-related protein B (SFTPB), and pulmonary surfactant-related protein C (SFTPC), while CLC_IL5RA-positive cardiomyocyte-like cells specifically express immune cell-related receptors, such as interleukin-5 receptor subunit α (IL5RA) and hematopoietic-specific transmembrane protein 4 (MS4A3).

[0393] Characterization of cell type-specific gene regulatory networks and pathways.

[0394] Next, we investigated cell-type-specific expression of surface and secretory protein-coding genes crucial for regulating cell-to-cell or cell-to-environment interactions. Most surface proteins (4,565 out of 5,480) and most secretory proteins (2,491 out of 2,933) were differentially expressed across 77 major cell types (FDR 0.05). For example, microglia specifically expressed sialic acid-binding immunoglobulin-like lectin 8 (SIGLEC8) and oxidized LDL endocytosis receptor (OLR1), both associated with Alzheimer's disease, while endothelial cells expressed roundabout-inducing receptor 4 (ROBO4) and endothelial cell adhesion molecule (ESAM), both involved in angiogenesis and vascular patterning. Similarly, different neurons were labeled with distinct cell surface transporters. For example, in the cerebellum, we observe the specific expression of the glycine neurotransmitter transporter SLC6A5 in inhibitory interneurons, the excitatory amino acid transporter SLC1A6 in Purkinje neurons, the potassium channel KCNK9 in granule neurons, and the sodium / potassium / calcium exchanger SLC24A4 in SLC24A4_PEX5L-positive inhibitory interneurons. Numerous similar examples exist of cell type-specific expression of secretory proteins. Particularly interesting is the unexpected cell type in the spleen (STC2_TLX1-positive cells) that specifically expresses the glycoprotein STC2, as well as TF TLX1 and NKX2-3, all of which are associated with mesenchymal precursors or stem cells.

[0395] Non-coding RNAs (NCRNAs) have been shown to play crucial roles in normal development and disease. In these data, 3,130 out of 10,695 NRNAs were differentially expressed across 77 major cell types (FDR 0.05). For example, ncRNAs were highly specific to microglial cells (RP11-489O18.1, RP11-480C22.1, RP11-10H3.1) or endothelial cells (AC011526.1, RP11-554D15.1, CTD-3179P9.1). While the biological significance of such cell type-specific ncRNAs remains unclear, it is noteworthy that their expression patterns were sufficient to separate the 77 major cell types into developmentally consistent groups.

[0396] The majority of transcription factors (TFs) were also expressed differentially across 77 major cell types (1,715 out of 1,984, FDR 0.05). Many of the most cell type-specific TFs were as expected, e.g., RBPJL in acinar cells, OLG1 and OLG2 in oligodendrocytes, and PAX7 in satellite cells. In other cases, cell type-specific TFs were noted as unexpected, such as stromal cell types characterized by the expression of lymphoid chemokines (CCL19_CCL21-positive cells), observed in the pancreas and specifically expressing TFs associated with immune activation.

[0397] The inventors attempted to directly predict TF target gene interactions via gene expression data. Briefly, candidate interactions were identified by the covariance between TF expression and target gene expression across the complete dataset. These interactions were further filtered by ChIP-seq linking and motif enrichment analysis ("Methods"). 56,272 candidate TF target gene links remained, including 706 TFs and 12,868 target genes. Of these 706 TF-binding gene sets, 220 showed enrichment of the corresponding TF (FDR 0.05) in manually clustered databases of TF networks (TRRUST) or Enrichr TF gene networks (e.g., the most enriched TRRUST TF for 330 genes binding to E2F1 was E2F1, with a regulatory p-value of 2.2e-14; the highest Enrichr TF for 1,219 genes binding to FLI1 was FLI1, with a regulatory p-value of 5.6e-122). When these 706 target genes assigned to TFs are rearranged and the analysis is repeated, none of the TF-binding gene sets are significantly enriched for their corresponding TFs at the same threshold.

[0398] Characterization of the development of blood systems across organs

[0399] The nature of this dataset provides an opportunity to investigate organ-specific differences in gene expression within a wide range of cell types, such as blood cells, endothelial cells, and epithelial cells. As a first such analysis, we reclustered 103,766 cells from all organs corresponding to hematopoietic cell types. We then performed Louvain clustering and further annotation of granular immune cell types based on published genetic markers. In some cases, we identified very rare cell types. For example, bone marrow cells are divided into microglia, macrophages, and diverse dendritic cell subtypes (CD1C+, S100A9+, CLEC9A+, and pDC). Microglia clusters originate mainly from the cerebrum and cerebellum and are well separated from macrophages, corresponding to their different developmental origins. Lymphoid cells were clustered into several groups, including B cells, NK cells, ILC3 cells, and T cells (the latter including thymic production orbitals). In addition, we recovered extremely rare cell types, including plasma cells (139 cells, representing 0.1% of total blood cells or 0.003% of the complete dataset, mostly from the placenta) and TRAF1+APC (189 cells, representing 0.2% of total blood cells or 0.005% of the complete dataset, mostly from the thymus and heart).

[0400] Gene expression markers for different immune cell types have been extensively studied, but these may be limited by definitions mediated through restricted sets of organ or cell types. In fact, we have found that many conventional immune cell markers are expressed in multiple cell types. For example, conventional markers for T cells were also expressed in macrophages and dendritic cells (CD4) or NK cells (CD8A), consistent with other studies. We calculated pan-organ cell type-specific markers across 14 blood cell types. For example, T cells specifically expressed CD8B and CD5 as expected, but also expressed TENM1. ILC3 cells, whose annotation is based on RORC and KIT expression, were more specifically labeled with SORCS1 and JMY. These and other pan-organ-defining markers may be useful in future studies for labeling and purifying human fetal blood cell types.

[0401] As expected, different organs exhibited remarkably different proportions of blood cells. For example, the liver contained the highest proportion of erythroblasts, consistent with its role as the primary site of fetal red blood cells, while T cells were concentrated in the thymus within the spleen and in B cells. Blood cells recovered from the cerebellum and cerebrum were almost entirely microglia. Collective analysis also allows for the identification of rare cell populations in specific organs. For example, we identified rare HSCs in the liver, spleen, and thymus, but also in the heart, lungs, adrenal glands, and intestines.

[0402] Focusing on erythrogenesis, we observed a continuous trajectory from HSCs to intermediate cell types, then to erythrocyte-basophil-megakaryocyte biased progenitor cells (EBMPs), which subsequently split into erythrocyte, basophilic, and megakaryocyte trajectories, consistent with recent studies of mouse fetal liver. This consistency was observed despite differences in species (human vs. mouse), technology (sci-RNA-seq3 vs. 10x), and organ (panorgan vs. fetal organ). Unsupervised clustering was performed, and terminology was adopted from the study to further divide the erythrocyte state continuum into three stages: early erythrocyte progenitor cells (EEP; labeled with SLC16A9 and FAM178B), delegated erythrocyte progenitor cells (CEP; labeled with KIF18B and KIF15), and terminal erythrocyte cells (ETD; labeled with TMCC2 and HBB). Early and late stages of megakaryocyte cells were also readily identified. The corresponding dynamics of genome-wide chromatin accessibility in erythrocyte lineages are further considered in the manual.

[0403] As expected, given the established roles in fetal erythrocytes, a significant proportion of immune cells in the liver and spleen corresponded to EEP, CP, and megakaryotic progenitor cells. Surprisingly, EEP, CEP, and megakaryotic progenitor cells were also observed in the adrenal glands in the nuclear samples studied. Trival contamination during adrenal retrieval is not a sufficient explanation, as these cell types, more common in the liver and spleen, were not observed. While confirmation by orthogonal analysis is needed, the results suggest the possibility of the adrenal glands as a site for fetal erythrocyte accretion.

[0404] Macrophages are even more broadly distributed. Next, all macrophages, along with microglia from the brain, were stained and independently subjected to UMAP visualization and Louvain clustering. Microglia were divided into three subclusters, one of which was labeled with IL1B and TNFRSF10D, likely representing active microglia involved in the inflammatory response. The other microglia clusters were labeled with the expression of TMEM119 and CX3CR1 (more common in the cerebrum) or PTPRC and CDC14B (more common in the cerebellum).

[0405] Macrophages outside the brain are clustered into three main groups: 1) antigen-presenting macrophages, mostly found in the GI tracheal organs (intestines and stomach), labeled by high expression of antigen-presenting (HLA DPB1, HLA DQA1) and inflammation-activating (AHR) genes; 2) perivascular macrophages, found in most organs, with specific expression of markers such as F13A1 and COLEC12, as well as novel markers such as RNASE1 and LYVE1; and 3) phagocyte macrophages, concentrated in the liver, spleen, and adrenal glands, with specific expression of markers such as CD5L, TIMD4, and VCAM1. Phagocyte macrophages are important for erythrocyte phagocytosis, and these observations in the adrenal glands are consistent with the aforementioned potential role as a site of fetal erythrocyte formation.

[0406] Characterization of endothelial and epithelial cells across organs

[0407] As a second analysis of single-cell types across multiple organs, the inventors reclustered cells from all organs corresponding to vascular endothelium, lymphoendothelium, or endocardium. These three groups readily separated from each other, and vascular endothelial cells were further clustered, at least to some extent, by organ. The organ-specific differences were more readily detectable than the differences between arteries, capillaries, and veins, and were consistent with previous cell atlases of adult mice.

[0408] Differential gene expression analysis identified 700 markers specifically expressed in a subset of endothelial cells (FDR 0.05, with more than a twofold difference in expression between the first and second clusters). For approximately one-third of these (236 out of 700) of the encoding membrane proteins, many appeared to correspond to potential special functions. For example, renal endothelial cells specifically expressed acid-detecting ion channel 2 (ASIC2), a mechanosensor involved in myogenic contraction and blood flow regulation within the kidney. Lung endothelial cells specifically expressed relaxin family peptide receptor 1 (RXFP1). RXFP1 specifically expressed sodium-dependent lysophosphatidylcholine transporter cotransporter 1 (MFSD2A), involved in endogenous nitric oxide-mediated vasodilation in the lungs, and MFSD2A is integrally involved in the establishment and function of the blood-brain barrier. Potential regulatory criteria for differential gene expression in endothelial subsets are discussed in the manual.

[0409] As a third analysis of the widely distributed cell types, epithelial cells derived from all organs were reclustered and subjected to UMAP visualization. Some epithelial cell types are highly organ-specific; for example, adenocarcinoma (pancreas) and alveolar cells (lung), and epithelial cells with similar functions, are generally clustered together. For instance, the expression programs of squamous epithelial cells (lung, stomach) are co-clustered with corneal and conjunctival epithelial cells (eye), and PDE1C_ACSM3-positive cells (stomach) are co-clustered with intestinal epithelial cells (intestine).

[0410] Within epithelial cells, two neuroendocrine cell clusters were identified. The simpler of these corresponded to adrenal chromaffin cells and were labeled with the specific expression of HHMX1 (NKX-5-3), a TF involved in the diversification of sympathetic neurons. The other cluster included neuroendocrine cells from multiple organs (stomach, intestine, pancreas, lung) and was labeled with the specific expression of NKX2-2, a TF that plays a crucial role in the differentiation of pancreatic islets and enterocrine cells. The inventors performed further analysis on the latter group and identified five subsets: 1) islet β cells labeled with insulin expression, 2) islet α / γ cells labeled with pancreatic polypeptide and glucagon expression, 3) islet δ cells labeled with somatostatin expression, 4) pulmonary neuroendocrine cells (PNECs) labeled with ASCL1, a TF that plays a crucial role in identifying this lineage in the lung, and 5) enterocrine cells. Enteroendocrine cells further comprise several subsets, including NEUROG-expressing islet ε-progenitor cells, TPH1-expressing chromaffin cells in both the stomach and intestines, and gastrin-expressing or cholecystokinin-expressing G / L / K / I cells. Finally, we observed ghrelin-expressing enteroendocrine progenitor cells in the stomach and intestines, as well as ghrelin-expressing endocrine cells in developing lungs. Because the diverse functions of neuroendocrine cells are closely linked to their secreted proteins, we identified 1,086 secreted protein-coding genes that are differentially expressed across neuroendocrine cells (FDR0.05). For example, PNEC showed specific expression of treoil factor 3, which is involved in mucosal protection and pulmonary calcification cell differentiation, gastrin-releasing peptide, which stimulates gastrin release from G cells in the stomach, and SCGB3A2, a surfactant associated with lung development.

[0411] As an exemplary example of how these data can be used to explore cell orbitals, we further investigated the diversification pathway of epithelial cells leading to renal tubular cells. We combined and reclustered ureteric bud metanephrocells and identified both progenitor and terminal renal epithelial cell types, and the differentiation pathway was remarkably consistent with recent studies of human fetal kidney. Differential gene expression analysis further evaluated the properties of TFs that potentially modulate their specifications. For example, nephron progenitor cells of the metanephrocellal orbital expressed high levels of mesenchymal and Meis homeobox genes (MEOX1, MEIS1, MEIS2), while podocytes specifically expressed MAFB and TCF21 / POD1. As another example, HNF4A was specifically expressed in proximal tubular cells, and mutations in this gene cause Fanconi tubular syndrome, a disease that specifically affects the proximal tubules. This has recently been shown to be necessary for proximal tubular formation in mice.

[0412] Comparison of Human and Mouse Developmental Atlases

[0413] To investigate developmental relationships between cell types, we then compared these data with our recent Mouse Organogenesis Cell Atlas (MOCA), which profiled 2 million cells from an entire embryo spanning the earlier mammalian developmental window, E9.5–E13.5.

[0414] As a first approach, the 77 major human cell types defined herein were compared to developmental trajectories defined by MOCA using the cell type cross-matching method described above. Briefly, this method uses non-negative least squares (NNLS) regression to select the best-matched cell type pairs from the two datasets. Most human cell types strongly match a single major mouse trajectory and sub-trajectories. These generally correspond to expectations and serve as a form of validation for both sets of annotations. Some discrepancies facilitated significant corrections to the MOCA annotations. Many human cell types and mouse trajectories lacking strong matches (summed NNLS regression coefficients < 0.6) corresponded to tissues excluded in the other dataset (e.g., mouse placenta, human skin, and gonads). Other ambiguities are likely due to gaps between developmental windows studied (e.g., adrenal cell types), rarity (e.g., bipolar cells), and / or complex relationships between cell types (e.g., fetal cell types derived from multiple embryonic trajectories).

[0415] As a second approach, we attempted to cluster human and mouse cells together. In short, we sampled 100,000 mouse embryonic cells (randomly) and 65,000 human fetal cells (up to 1,000 cells from each of the 77 cell types) from MOCA and subjected them to Seurat's recently described strategy to integrate the cross-species scRNA-seq datasets. The distribution of mouse cells in the resulting UMAP-based visualizations was remarkably similar to our global analysis of MOCA. Furthermore, and surprisingly, the cells were distributed in a generally reasonable manner with respect to both developmental and temporal relationships, rather than just spatial organ locations. For example, we observed that human fetal endothelial cells, hematopoietic cells, hepatocytes, epithelial cells, and mesenchymal cells all mapped to their corresponding mouse embryonic orbitals. Human fetal brain neurons and cerebellar neurons overlapped with mouse embryonic neural tube orbitals, but human fetal neural crest derivatives, such as ENS neurons, visceral neurons, sympathetic neuroblasts, and chromaffin cells, clustered separately from their corresponding mouse embryonic orbitals, likely due to excessive differences between species or developmental stages. As expected, human ENS glia and Schwann cells overlapped with mouse embryonic PNS gira suborbitals. Human fetal astrocytes clustered with mouse embryonic neuroepithelial orbitals (mouse astrocytes do not develop until E18.5). Human fetal oligodendrocytes overlapped with rare mouse embryonic suborbitals (Pdgfra+ glia), which, upon reflection, correspond to oligodendrocyte precursor cells (OPCs; Olig1+, Olig2+, Brinp3+), casting doubt on previous annotations of different Oligo1+ suborbitals as oligodendrocyte precursors.

[0416] To visualize the more detailed relationships between human fetal cells and mouse embryonic cells, a similar integrated analysis strategy was applied to extract human and mouse cells from hematopoietic, endothelial, and epithelial orbitals. Data from this fetal human cell atlas allows for the easy deconvolution of “whole embryo” mouse data into granular functional or spatial groups. For example, subsets of the mouse “leukocyte” orbital map to specific human blood cell types, such as HSCs, microglia, macrophages (liver and spleen), macrophages (other organs), and DCs. These subsets were further demonstrated by the expression of relevant blood cell markers. Similarly, the inventors observed that relevant subsets of mouse / human endothelial and epithelial cells map to each other. This approach may be useful for obtaining gene expression programs of specific strains of progenitor cells at developmental points where access or anatomical degradation is difficult. For instance, within mouse cells previously labeled as foregut epithelial orbitals, it is possible to degrade factors likely to be used for gastric versus pancreatic development.

[0417] Consideration

[0418] The successful development of a functional human fetus is a remarkable process characterized by cell proliferation and differentiation processes across three major developmental stages.

[0419] Following a short embryonic period (two weeks from fertilization) involving simple cell proliferation and implantation in the uterus, the embryogenetic stage continues with gastrulation, neurogenesis, and organogenesis, characterized by vigorous cell differentiation and the generation of organ precursors. By the end of the 10th week of gestation, the embryo has acquired the basic form known as the fetus. Over the next 20 weeks, various organs continue to grow and mature, and diverse terminally differentiated cell types are generated from precursors.

[0420] Both the embryonic and embryogenetic stages have been intensively profiled at single-cell resolution in humans or model systems (i.e., mice) using a shared early developmental program. The later developmental stages (fetal stages) exhibit different developmental programs and durations in Homo sapiens and other species. Furthermore, due to the greater complexity of organs and technical limitations, obtaining a comprehensive picture of cell dynamics at this stage is difficult. While several single-cell studies of fetal development have recently been published, most are limited to specific organs or cell lineages and fail to provide a complete picture of organ development.

[0421] Materials and methods:

[0422] Mammalian cell culture and nuclear extraction

[0423] All mammalian cells were cultured at 5% CO2 and 37°C, and maintained in high-glucose DMEM (Gibco catalog no. 11965) supplemented with 10% FBS and 1X Pen / Strep (Gibco catalog no. 15140122; 100 U / mL penicillin, 100 μg / mL streptomycin). Cells were trypsinized with 0.25% trypsin-EDTA (Gibco catalog no. 25200-056) and divided into 1:10 portions three times a week.

[0424] All cell lines were trypsinized, spun down at 300xg for 5 minutes (4°C), and washed once with 1X ice-cold PBS. 5M cells were combined and lysed using 1 mL of ice-cold cell lysis buffer (modified to contain 10 mM Tris-HCl, pH 7.4, 10 mM NaCl, 3 mM MgCl2, and 0.1% IGEPAL CA-630, 1% SUPERase InRNase inhibitor, and 1% BSA). Filtered nuclei were then transferred to a new 15 mL Falcon tube, pelletized by centrifugation at 500xg at 4°C for 5 minutes, and washed once with 1 mL of ice-cold cell lysis buffer. The nuclei were fixed on ice for 15 minutes in 4 mL of ice-cold 4% paraformaldehyde (EMS). After fixation, the nuclei were washed twice in 1 mL of nuclear washing buffer (IGEPAL-free cell lysis buffer) and resuspended in 500 μL of nuclear washing buffer. The sample was placed in 100 μL of each tube, divided into five tubes, and rapidly frozen in liquid nitrogen.

[0425] Preparation and nuclear extraction of human fetal tissue

[0426] Human fetal tissues were processed together to reduce batch effects. Each organ was pulverized into tissue powder using a hammer (on dry ice) and mixed before sampling. First, 1 mL of ice-cold cell lysis buffer (10 mM Tris-HCl, pH 7.4, 10 mM NaCl, 3 mM MgCl2, and 0.1% IGEPAL CA-630) was used. 530.1–1 g of powder was incubated with (modified to also contain 1% SUPERase In and 1% BSA), and then transferred to a 40 μm cell strainer (Falcon). The tissue was homogenized in 4 mL of cell lysis buffer using a syringe plunger (5 mL, BD) with a rubber tip. The filtered nuclei were then transferred to a new 15 mL tube (Falcon), pelletized by centrifugation at 500 x g for 5 minutes, and washed once with 1 mL of cell lysis buffer. The nuclei were fixed in 5 mL of ice-cold 4% paraformaldehyde (EMS) for 15 minutes on ice. After fixation, the nuclei were washed twice in 1 mL of nuclear washing buffer (cell lysis buffer without IGEPAL) and resuspended in 500 μL of nuclear washing buffer. 250 μL of the sample was placed in each tube, divided into two tubes, and rapidly frozen in liquid nitrogen. In the case of human cell extraction and paraformaldehyde fixation in certain organs (kidney, pancreas, intestine, and stomach).

[0427] Preparation and sequencing of sci-RNA-seq3 libraries

[0428] Paraformaldehyde-fixed nuclei were processed with minor modifications, similar to the published sci-RNA-seq3 protocol. Briefly, thawed nuclei were permeabilized on ice for 3 minutes using 0.2% TritonX-100 (in nuclear wash buffer), followed by brief sonication (Diagenode, low power mode, 12 seconds) to reduce nuclear aggregation. The nuclei were then washed once with nuclear wash buffer and filtered through a 1 mL Flowmi cell strainer (Flowmi). The filtered nuclei were spun down at 500xg for 5 minutes and resuspended in nuclear wash buffer. The nuclei from each sample were then distributed into multiple individual wells in four 96-well plates. The link between well IDs and mouse embryos was recorded for downstream data processing. In each well, 80,000 nuclei (16 μL) were mixed with 8 μL of 25 μM immobilized oligo-dT primer ((5'- / 5Phos / CAGAGCNNNNNNNN[10bp barcode]TTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTT-3' (SEQ ID NO: 1), where "N" is any base; IDT) and 2 μL of 10 mM dNTP mix (Thermo), denatured at 55°C for 5 minutes, and immediately placed on ice. Then, 8 μL of 5X Superscript IV First-Strand Buffer (Invitrogen), 2 μL of 100 mM DTT (Invitrogen), 2 μL of SuperScript IV reverse transcriptase (200 U / μL, Invitrogen), and 2 μL of RNaseOUT Recombinant Ribonuclease 14 μL of the first-chain reaction mix containing the inhibitor (Invitrogen) was added to each well. Reverse transfer was performed by incubating the plate at a gradient temperature (4°C for 2 minutes, 10°C for 2 minutes, 20°C for 2 minutes, 30°C for 2 minutes, 40°C for 2 minutes, 50°C for 2 minutes, and 55°C for 10 minutes).

[0429] After the reverse transcription reaction, 60 μL of nuclear dilution buffer (10 mM Tris-HCl, pH 7.4, 10 mM NaCl, 3 mM MgCl2, and 1% BSA) was added to each well. The nuclei from all wells were pooled together and spun down at 500xg for 10 minutes. The nuclei were then resuspended in nuclear wash buffer and redistributed into four separate 96-well plates, each containing 20 μL of Quick ligase buffer (NEB), 2 μL of Quick DNA ligase (NEB), 10 μL of nuclei in nuclear wash buffer, and 8 μL of barcoded ligation adapter (100 μM, 5'-GCTCTG [barcode A for 9 bp or 10 bp] / dideoxy U / ACGACGCTCTTCCGATCT [reverse complement of barcode A]-3' (SEQ ID NO: 2)) in each well. Ligation was performed at 25°C for 10 minutes. After the ligation reaction, 60 μL of nuclear dilution buffer (10 mM Tris-HCl, pH 7.4, 10 mM NaCl, 3 mM MgCl2, and 1% BSA) was added to each well. The nuclei from all wells were pooled together and spun down at 600xg for 10 minutes.

[0430] The nuclei were washed once with nuclear wash buffer, filtered once through a 1 mL Flowmi cell strainer (Flowmi), counted, and distributed into eight 96-well plates, each containing 2,500 cells in 5 μL of nuclear wash buffer and 3 μL of elution buffer (Qiagen). Then, 1.33 μL of mRNA second-chain synthesis buffer (NEB) and 0.66 μL of mRNA second-chain synthase (NEB) were added to each well, and second-chain synthesis was carried out at 16°C for 180 minutes.

[0431] For tagging, each well was mixed with 11 μL of Nextera TD buffer (Illumina) and 1 μL of i7-only TDE1 enzyme (62.5 nM, Illumina, diluted in Nextera TD buffer (Illumina)), and then incubated at 55°C for 5 minutes to tag. The reaction was then stopped by adding 24 μL of DNA binding buffer (Zymo) per well and incubated at room temperature for 5 minutes. Each well was then purified using 1.5x AMPure XP beads (Beckman Coulter). For the elution step, 8 μL of nuclease-free water, 1 μL of 10X USER buffer (NEB), and 1 μL of USER enzyme (NEB) were added to each well and incubated at 37°C for 15 minutes. Another 6.5 μL of elution buffer was added to each well. The AMPure XP beads were removed using a magnetic stand, and the eluted product (16 μL) was transferred to a new 96-well plate.

[0432] For PCR amplification, each well (16 μL of product) was mixed with 2 μL of 10 μM indexed P5 primer (5'-AATGATACGGCGACCACCGAGATCTACAC[i5]ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3' (SEQ ID NO: 3); IDT), 2 μL of 10 μM P7 primer (5'-CAAGCAGAAGACGGCATACGAGAT[i7]GTCTCGTGGGCTCGG-3' (SEQ ID NO: 4), IDT), and 20 μL of NEBNext High-Fidelity 2x PCR MASTER Mix (NEB). Amplification was performed using a program of 5 minutes at 72°C, 30 seconds at 98°C, or 10 seconds at 98°C, 30 seconds at 66°C, 1 minute at 72°C for 12-16 cycles, and finally 5 minutes at 72°C.

[0433] After PCR, the samples were pooled and purified using 0.8 volt AMPure XP beads. Library concentrations were determined using Qubit (Invitrogen), and the libraries were visualized by electrophoresis on a 6% TBE-PAGE gel. All libraries were sequenced on a single NovaSeq platform (Illumina) (read 1: 34 cycles, read 2: 52 cycles, index 1: 10 cycles, index 2: 10 cycles).

[0434] Paraformaldehyde-fixed cells were treated with slight modifications, similar to the fixed nuclei, as follows: Frozen-fixed cells were thawed in a 37°C water bath, spun down at 500xg for 5 minutes, and incubated on ice for 3 minutes in 500 μL of PBSR (1x PBS, pH 7.4, 1% BSA, 1% SuperRnaseIn, 1% 10 mM DTT) containing 0.2% Triton X-100. The cells were pelleted and resuspended in 500 μL of nuclease-free water containing 1% SuperRnaseIn. 3 mL of 0.1 N HCl was added to the cells for incubation on ice for 5 minutes (7). To neutralize the HCl, 3.5 mL of Tris-HCl (pH=8.0) and 35 μL of 10% Triton X-100 were added to the cells. The cells were pelleted and washed with 1 mL of PBSR. The cells were pelleted and resuspended in 100 μL of PBSI (1x PBS, pH 7.4, 1% BSA, 1% SuperRnaseIn). The subsequent steps were similar to the sci-RNA-seq3 protocol described above (using paraformaldehyde-fixed nuclei), but with slight modifications: (1) 20,000 fixed cells were distributed per well (instead of 80,000 nuclei) for reverse transcription; (2) all nuclear wash buffers were replaced with PBSI in subsequent steps; and (3) all nuclear dilution buffers were replaced with PBS + 1% BSA.

[0435] Read sequencing process

[0436] We performed single-cell RNA-seq read alignment and gene count matrix generation using a pipeline developed for sci-RNA-seq3 with some modifications. Specifically, we converted base calls to fastq format using Illumina's bcl2fastq / v2.16 and demultiplexed them based on PCR i5 and i7 barcodes using the maximum likelihood demultiplexing package deML with default settings. Downstream sequencing and single-cell digital expression matrix generation were similar to sci-RNA-seq, except that the RT index was combined with a hairpin adapter index; therefore, mapped reads were split into constituent cell indices by demultiplexing the reads using both the RT index and ligation index (ED<2, including insertions and deletions). In short, the demultiplexed reads were filtered based on the RT index and ligation index (ED<2, including insertions and deletions), and the adapters were clipped using trim_galore / v0.4.1 with default settings. Modified reads were mapped to the human reference genome (hg19) of human fetal nuclei, or the chimeric reference genome of human hg19, and to mouse mm10 with HEK293T and NIH / 3T3 mixed nuclei, using STAR / v 2.5.2b with default settings and gene annotation (GENCODE V19 for humans, GENCODE VM11 for mice). Uniquely mapped reads were extracted, and duplicates were removed using unique molecular identifier (UMI) sequences (ED<2, including insertions and deletions), reverse transcription (RT) index, hairpi...

Claims

1. A method for identifying subpopulations of cells that possess biological characteristics, (a) To provide a single-cell sequencing library, The sequencing library includes multiple modified target nucleic acids, The modified target nucleic acid includes at least one index sequence, (b) Examining the sequencing library by targeted sequencing in order to identify index sequences related to biological characteristics, The aforementioned biological features are present in a subpopulation of cells in the single-cell sequencing library but not in other cells in the library, and the identified index sequence is present in the same subpopulation of cells as the subpopulation of cells in which the biological features are present. The index sequence related to the aforementioned biological characteristics is a marker index sequence, (c) Modifying the sequencing library to obtain a sublibrary, The sublibrary includes an enhanced representation of the modified target nucleic acid, including the marker index sequence, compared to other modified target nucleic acids present in the sequencing library that do not include the marker index sequence. (d) A method comprising determining the nucleotide sequence of the modified target nucleic acid, which includes a marker index sequence.

2. The method according to claim 1, wherein the single-cell sequencing library comprises nucleic acids from multiple samples.

3. The method according to claim 2, wherein the plurality of samples include (i) a sample of the same tissue obtained from different organisms, (ii) a sample of different tissues from one organism, or (iii) a sample of different tissues from different organisms.

4. The method according to claim 1, wherein in step (b), two or more marker index sequences are identified.

5. The method according to claim 1, wherein the single-cell sequencing library includes a target nucleic acid representing the whole genome or a subset of the genome of the cell or nucleus.

6. The method according to claim 5, wherein the subset of the genome includes target nucleic acids representing the cell or nucleus transcriptome, accessible chromatin, DNA, structural state, or protein.

7. The method according to any one of claims 1 to 6, wherein the modification includes enriching the modified target nucleic acid containing the marker index sequence.

8. The method according to claim 7, wherein the concentration includes a hybridization-based method.

9. The method according to claim 8, wherein the hybridization-based method comprises hybrid capture, amplification, or CRISPR(d)Cas9.

10. The method according to claim 9, wherein the modification includes depleting the modified target nucleic acid which does not include the marker index sequence.

11. The method according to claim 10, wherein the depletion includes a hybridization-based method.

12. The hybridization-based method according to claim 11, comprising hybrid capture, amplification, or CRISPR(d)Cas9.

13. The method according to claim 1, wherein the biological characteristics include a nucleotide sequence indicating the type of species.

14. The method according to claim 13, wherein the type of the species includes the species of the cell.

15. The method according to claim 14, wherein the biological features include a nucleotide of the 16s subunit, the 18s subunit, or the ITS non-transcribed region.

16. The method according to claim 1, wherein the aforementioned biological features include a nucleotide sequence indicating a cell class.

17. The method according to claim 16, wherein the cell class includes an expression pattern, an epigenetic pattern, immunorecombination, or a combination thereof.

18. The method according to claim 17, wherein the epigenetic pattern comprises a methylation label, a methylation pattern, accessible DNA, or a combination thereof.

19. The method according to claim 1, wherein the biological features include a nucleotide sequence indicating a disease state or risk.

20. The method according to claim 19, wherein the disease state or risk includes a mutated DNA sequence, a mutation expression pattern, or a mutation epigenetic pattern correlated with the disease.

21. The method according to claim 20, wherein the mutated DNA sequence comprises at least one single nucleotide polymorphism.

22. The method according to claim 21, wherein the mutation expression pattern includes the expression of a biomarker.

23. The method according to claim 22, wherein the mutational epigenetic pattern includes a methylation label and a methylation pattern.

24. The method according to claim 1, wherein the modified target nucleic acid comprises consecutive indices of at least two compartment-specific index sequences, and there are no seven or more nucleotides between the two index sequences.

25. The method according to claim 24, wherein the continuous index is present at each end of the modified target nucleic acid.

26. The method according to claim 24 or 25, wherein the length of the continuous index is at least 55 nucleotides.

27. The method according to any one of claims 24 to 26, wherein one copy of the continuous index is present in the modified target nucleic acid.

28. The method according to any one of claims 24 to 26, wherein the two copies of the continuous index are present in the modified target nucleic acid.

29. The method according to claim 1, wherein the plurality of modified target nucleic acids in the sequencing library represent at least 100,000 different cells or nuclei.

30. Providing the aforementioned single-cell sequencing library is, The method according to claim 1, comprising processing a sample to prepare a library, wherein the sample is a metagenomics sample obtained from an organism.

31. The method according to claim 30, wherein the organism is a mammal.

32. The method according to claim 30 or 31, wherein the metagenomics sample includes tissue suspected of containing symbiotic or pathogenic microorganisms.

33. The method according to claim 32, wherein the microorganism is a prokaryote or a eukaryote.

34. The method according to any one of claims 30, 31, or 33, wherein the metagenomics sample includes a microbiome sample.

35. Providing the aforementioned single-cell sequencing library is, The method according to claim 1, comprising processing a sample to prepare a library, wherein the sample is from a living organism.

36. The method according to claim 35, wherein the organism is a mammal.

37. The method according to claim 35, wherein the primary source of nucleic acids from the sample includes RNA.

38. The method according to claim 37, wherein the RNA includes mRNA.

39. The method according to claim 35, wherein the primary source of nucleic acids from the sample includes DNA.

40. The method according to claim 39, wherein the DNA includes whole cell genomic DNA.

41. The method according to claim 40, wherein the whole cell genomic DNA includes nucleosomes.

42. The method according to claim 35, wherein the primary source of nucleic acids from the sample includes cell-free DNA.

43. The method according to claim 35, wherein the sample contains cancer cells.

44. A method for sequencing single cells or single nuclei, (a) Uniquely indexing the nucleic acids of each cell or nucleus in the sample, thereby creating an indexed library of each cell or nucleus, (b) Identifying one or more indexed libraries of interest from step (a) using biological features, wherein the biological features are present in a subpopulation of cells or nuclei of the sample and not in other cells of the sample. (c) Concentrating the target indexed library in step (b), thereby producing a concentrated library, A method comprising (d) sequencing the concentrated library from step (c).

45. The method according to claim 44, wherein identifying the target indexed library includes sequencing the index.

46. The method according to claim 44 or 45, wherein the library is derived from DNA, RNA, or proteins of the cell or the nucleus.

47. The method according to any one of claims 44 to 46, wherein the biological feature is DNA, RNA, or protein, or a combination thereof.

48. The method according to any one of claims 44 to 46, wherein the unique indexing in step (a) comprises associating at least two different indices with the nucleic acids of the cell or the nucleus.

49. The method according to claim 48, wherein the at least two different indices are consecutive indices.

50. The method according to any one of claims 44 to 46, wherein the concentrated library is prepared by positive concentration.

51. The method according to claim 50, wherein the positive enrichment includes amplification.

52. The method according to claim 50, wherein the positive concentration includes a scavenging agent.

53. The method according to claim 50, wherein the positive enrichment includes a solid support.

54. The method according to claim 50, wherein the concentrated library is prepared by negative concentration.

55. A method for sequencing single cells or single nuclei, (a) To provide a sample, wherein the sample contains a plurality of nuclei or cells, (b) Associating a first index with each nucleus or cell in the sample, (c) Dividing the sample into multiple sections, (d) Associating a second index with each nucleus or cell of the plurality of compartments, (e) Pooling the multiple partitions, (f) Sequencing the pooled partitions, (g) Identifying a combination of a first index and a second index associated with a biological feature, wherein the biological feature is present in a subpopulation of cells or nuclei of the sample but not in other cells of the sample. A method comprising (h) enriching the biological features from the pooled compartments using the identified combination of the first index and the second index from step (g).